HTX301
The first reference chip on the HyperThought™ platform to run ultra-large models on a single PCIe card.700B-parameter LLM inference, on-prem, at about 240W — no GPU cluster required.

The idea
A different architecture for LLMs
Ultra-large models don't need a superscalar GPU cluster. HTX301 disaggregates prefill and decode workloads and prioritizes a decode-first silicon design, using the LISA™ instruction set to scale from 4B to 700B parameters on the same architecture.
The result eliminates massive GPU clusters, NVLink/NVSwitch interconnects, and complex liquid cooling — delivering data sovereignty, predictable cost, and deterministic performance for agentic AI.
"The era of needing superscalar GPU clusters for ultra-large LLMs is over."— William Wei, CMO
No cluster tax
Skip NVLink/NVSwitch, multi-node orchestration, and the cooling that comes with them.
Sovereign by default
Inference never leaves your premises — data, weights, and prompts stay in-house.
Scale without waste
One architecture from 4B to 700B; provision for the model you run, not the worst case.
Specifications
HTX301 at a glance


Figures reflect the HTX301 reference platform and are subject to change ahead of general availability.
Early access
Request access to HTX301
Built for ultra-large model inference with enterprise-grade privacy, power efficiency, and deployment simplicity. Register below and our team will follow up with specifications, evaluation details, and next steps.
- 700B-parameter inference on a single PCIe card
- Decode-first acceleration · unified prefill/decode orchestration
- On-prem deployment with full data sovereignty
Prefer email? sales@skymizer.ai