Skymizer Announces HTX301 — Reinventing On-Prem AI Inference
HTX301 enables 700B-parameter LLM inference locally at just ~240W on a single PCIe card — no GPU cluster required.

Skymizer today announced HTX301, the first reference chip built on the HyperThought™ platform — redefining how enterprises deploy and scale AI inference.
For the first time, ultra-large models can run on a single PCIe card. Powered by six HTX301 chips and 384GB of memory, enterprises can now execute 700B-parameter LLM inference locally at just ~240W — eliminating the need for massive GPU clusters, NVLink/NVSwitch interconnects, and complex cooling infrastructure.
Built for the new era of inference-dominant AI, HyperThought™ introduces a fundamentally different approach. By disaggregating prefill and decode workloads and pairing decode-first silicon with an intelligent software orchestration stack, HTX301 enables higher utilization, lower latency, and significantly improved power efficiency across real-world deployments.
HyperThought™ scales seamlessly from on-device to on-prem environments under a unified architecture powered by LISA™ (Language Instruction Set Architecture), allowing enterprises to right-size deployments from 4B to 700B models without over-provisioning.
The result is a new class of AI infrastructure: data sovereignty, predictable cost, and deterministic performance — unlocking agentic AI workflows across enterprise applications without the hidden tax of per-token cloud inference.
“Inference has become the dominant AI workload, and infrastructure needs to reflect that reality. The era of needing superscalar GPU clusters for ultra-large LLMs is over. HyperThought shifts AI from hyperscaler-only complexity to single-card simplicity for every enterprise.”
— William Wei, Chief Marketing Officer, Skymizer
“Purpose-built decode hardware paired with an intelligent software stack that orchestrates every inference workload — that’s how you disaggregate P/D at scale.”
— Luba Tang, Chief Technology Officer, Skymizer
As AI models scale from billions to trillions of parameters, HTX301 marks a decisive step beyond brute-force GPU scaling — delivering a simpler, more efficient, and enterprise-ready path to AI deployment.