Silicon · Early access

HTX301

The first reference chip on the HyperThought™ platform to run ultra-large models on a single PCIe card.700B-parameter LLM inference, on-prem, at about 240W — no GPU cluster required.

HTX301 evaluation board
700Bparameters, on a single card
~240Wtotal board power
384GBon-card memory

The idea

A different architecture for LLMs

Ultra-large models don't need a superscalar GPU cluster. HTX301 disaggregates prefill and decode workloads and prioritizes a decode-first silicon design, using the LISA™ instruction set to scale from 4B to 700B parameters on the same architecture.

The result eliminates massive GPU clusters, NVLink/NVSwitch interconnects, and complex liquid cooling — delivering data sovereignty, predictable cost, and deterministic performance for agentic AI.

"The era of needing superscalar GPU clusters for ultra-large LLMs is over."— William Wei, CMO

No cluster tax

Skip NVLink/NVSwitch, multi-node orchestration, and the cooling that comes with them.

Sovereign by default

Inference never leaves your premises — data, weights, and prompts stay in-house.

Scale without waste

One architecture from 4B to 700B; provision for the model you run, not the worst case.

Specifications

HTX301 at a glance

Configuration6× HTX301 LPU chips on a single PCIe card
Memory384 GB (standard LPDDR-class)
Model reach4B → 700B parameters, no over-provisioning
Power≈ 240 W total
ProcessT28nm mature node
ArchitectureLISA™ ISA · prefill/decode disaggregation · decode-first
Form factorPCIe card — on-prem deployment
TargetEnterprise & agentic AI inference
HTX301 EVB front
Evaluation board — front
HTX301 EVB back
Evaluation board — back

Figures reflect the HTX301 reference platform and are subject to change ahead of general availability.

Early access

Request access to HTX301

Built for ultra-large model inference with enterprise-grade privacy, power efficiency, and deployment simplicity. Register below and our team will follow up with specifications, evaluation details, and next steps.

  • 700B-parameter inference on a single PCIe card
  • Decode-first acceleration · unified prefill/decode orchestration
  • On-prem deployment with full data sovereignty

Prefer email? sales@skymizer.ai

We'll only use your details to respond to this request. See our Privacy Policy.