Free to read. Sign up to save tools and get alerts when they change. Plus 900+ more AI tool profiles.

Sign up free
6 min read·Updated August 19, 2026

Etched Sohu

Etched logoBy Etched

Etched Sohu is an inference chip purpose-built for the transformer architecture — an application-specific integrated circuit that hard-wires transformer inference on TSMC's 4-nanometer process for far higher speed and lower cost than general-purpose GPUs. Etched has booked more than $1 billion in orders and raised $700 million led by Jane Street in August 2026 at a $21 billion valuation, double its July round; Jane Street is also running Etched's first shipped cluster in its own data center.

Share

Listen to this overview

Free preview · first 0:30
0:00 / 0:30

Unlock audio and more

Audio streaming, downloadable PDFs and certificates come with Plus and Pro.

Learning Objectives

  • Understand what makes Sohu different from a general-purpose GPU
  • Explain why specializing a chip for the transformer architecture can improve inference speed and cost
  • Identify where Etched fits among AI inference challengers to Nvidia

What Is Etched Sohu?

Etched Sohu is an AI inference chip built by Etched, an American semiconductor startup founded in 2022 by Harvard dropouts Gavin Uberti, Robert Wachen, and Chris Zhu. Where most AI chips — including Nvidia's GPUs — are general-purpose processors that can run any kind of model, Sohu makes a deliberate bet in the opposite direction: it is designed to run one thing extremely well.

That one thing is the transformer, the neural-network architecture behind virtually every modern large language model, from GPT and Claude to Gemini and Llama. Sohu is an application-specific integrated circuit (ASIC) — a chip whose logic is hard-wired for a fixed task rather than programmable for many. Etched hard-wires the matrix-multiplication patterns specific to transformer inference directly into silicon, fabricated on TSMC's 4-nanometer process.

💡Key Concept

Specialization versus flexibility: A GPU is a Swiss-army knife — it can train and run any model, but it spends transistors and energy on that flexibility. An ASIC like Sohu is a single-purpose tool: it can only run transformers, but because it does nothing else, far more of the chip is dedicated to the actual work. The bet is that the transformer has won decisively enough that specializing for it is worth giving up the flexibility.

Why a Transformer-Only Chip?

The economics of AI have shifted. Training a frontier model is a one-time cost; inference — actually running the model to answer queries — happens billions of times and now dominates the ongoing cost of operating AI at scale. That makes inference efficiency the battleground.

Etched's thesis is that once an architecture becomes as dominant as the transformer, the industry can afford to bake it into hardware. By removing the general-purpose overhead of a GPU, Sohu aims to deliver substantially more throughput per dollar and per watt on transformer inference specifically. Etched sells both the chips and full frontier inference clusters — turnkey systems that pair Sohu chips with custom racks and software so a customer can deploy inference capacity without assembling it piece by piece.

Etched's newer designs push specialization one level deeper by splitting inference into its two natural phases. A low-voltage prefill chip digests the incoming prompt, while a cluster-scale memory interconnect accelerates the decode phase, where the model generates output one token at a time. Because prompt-processing and token-generation stress hardware very differently, handling each with purpose-built silicon is the same specialization bet applied to the inference pipeline itself.

⚠️Warning

The risk of specialization: A transformer-only chip is a bet on the transformer staying dominant. If a fundamentally different architecture displaces it, a general-purpose GPU can adapt where a hard-wired ASIC cannot. Etched is wagering that the transformer's lead is durable enough to make that risk worth taking.

Traction and Backing

Etched has booked more than $1 billion in orders for its inference systems and is one of the most closely watched challengers in AI hardware. In July 2026 it raised a $300 million Series C led by Sequoia, with Andreessen Horowitz and memory-maker SK Hynix joining — a round that more than doubled the company's valuation from six months earlier, to $10.3 billion.

One month later, on August 18, 2026, it raised $700 million more, led by Jane Street, at a $21 billion valuation — doubling again in roughly four weeks. The unusual part is that Jane Street is not only leading the round but running the product: the quantitative-trading firm has installed Etched's first shipped cluster in its own data center, saying it is "pleased with the early results." A lead investor who is also the first production customer is a materially stronger signal than a valuation on its own, because the money and the deployment are being risked by the same party.

⚠️Warning

Doubling a valuation in a month says as much about the market as the company. Etched went from $10.3 billion to $21 billion between mid-July and mid-August 2026 on the strength of one installed system. That is a real milestone — silicon that ships beats silicon that benchmarks — but a single named customer is not yet evidence of broad demand, and specialized inference chips have a long history of impressive first deployments that never generalized. Treat the valuation as a statement about how badly the market wants a credible alternative to Nvidia, not as a measure of shipped volume.

Earlier backers include quantitative-trading firms Jane Street, Hudson River Trading, and Two Sigma, along with Ribbit Capital, and an angel roster that reads like a who's-who of AI: Andrej Karpathy, Geoffrey Hinton, and Peter Thiel among them.

It competes with Nvidia's inference GPUs as well as other specialized challengers such as Cerebras and Groq, and with the custom in-house chips being built by Amazon, Google, Microsoft, and OpenAI.

Pricing

Sohu chipsCustom / enterprise
  • Direct hardware purchase
  • Volume-based pricing
Frontier inference clustersCustom / enterprise
  • Turnkey chip, rack, and software systems
  • Deployment and integration support

Etched sells to enterprises and AI infrastructure operators; there is no self-serve or consumer pricing. Access is arranged directly through the company.

  • Cerebras Inference — wafer-scale inference challenger with a different specialization approach
  • Groq Cloud — low-latency inference on a purpose-built LPU

Strengths

  • Purpose-built for the dominant workload — specializing for the transformer targets exactly where inference spending concentrates
  • Throughput and cost focus — the design goal is more transformer inference per dollar and per watt than a general-purpose GPU
  • Turnkey systems — frontier inference clusters let customers buy deployable capacity, not just chips
  • Strong backing and real demand — more than $1 billion booked and a top-tier investor and angel roster

Limitations and Considerations

  • Transformer-only — Sohu cannot run non-transformer models, so it is a bet on the architecture's continued dominance
  • Enterprise-only — no self-serve access; relevant to infrastructure operators, not individual developers
  • Young company — Etched is scaling from orders to at-volume delivery, and execution risk remains

Key Takeaways

  • Etched Sohu is an inference ASIC purpose-built for the transformer architecture, fabricated on TSMC's 4-nanometer process
  • Its thesis is that inference now dominates AI costs, so specializing hardware for the dominant architecture beats general-purpose flexibility
  • Etched has booked more than $1 billion in orders and doubled its valuation twice in two months — a July 2026 Series C led by Sequoia took it to $10.3 billion, and a $700 million round led by Jane Street on August 18 took it to $21 billion
  • Jane Street both led that round and installed Etched's first shipped cluster in its own data center, which is a stronger signal than the valuation itself — though one named customer is not yet evidence of broad demand
  • Its newer designs split inference into a purpose-built prefill chip for the prompt and a cluster-scale interconnect for the decode phase — specialization applied to the inference pipeline itself
  • The trade-off is flexibility: a transformer-only chip wins only as long as the transformer stays dominant

Keep track of the tools you’re evaluating

  • The AI Hub on a phone: a 12-day AI Skill Streak and an expanded Content updates alert listing the saved items that changed.
  • Recommended for you on a phone: nine personalised suggestions labelled Trending in AI news, On your saved list, and Popular.
  • My AI Tools on a phone: saved tools including GitHub Copilot and OpenAI Codex, each with an Updated badge.

Swipe for Recommended for you and My AI Tools

Your AI Hub — sample data.

📰Etched Sohu in the News

Showing the 3 stories where Etched Sohu is tagged in Top AI Stories.

Other tools in AI Chips & Hardware (12 of 22)

Show 7 more →
🧭Recommended for you

Optional detours — these connect to what you just read, and your next lesson will be waiting.