📋About Fireworks AI
Updated August 26, 2026Fireworks AI is an inference and fine-tuning platform that helps companies turn open-source foundation models into specialized, production-grade systems trained on their own data. Founded in 2022 by a team from Meta's PyTorch group and led by CEO Lin Qiao — formerly the Senior Director of Engineering who ran PyTorch at Meta — the company is based in Redwood City, California. Its stack centers on proprietary serving technology: the FireAttention inference kernel, speculative decoding, and the FireOptimizer adaptive engine, which together push high throughput at low latency. Fireworks serves a broad catalog of open models — including Llama, DeepSeek, Qwen, and Mixtral — through an OpenAI-compatible API, and now handles more than 40 trillion tokens a day for customers such as Cursor and Harvey. In July 2026 it raised $1.5 billion in a Series D led by Atreides Management, Index Ventures, and TCV, with backing from Nvidia.
🛠️Products & Tools (1)
Inference and fine-tuning platform for serving specialized open-source models at production scale — powered by the FireAttention kernel, speculative decoding, and the FireOptimizer adaptive engine. OpenAI-compatible API across Llama, DeepSeek, Qwen, and Mixtral.
Keep track of the companies you’re watching
- Save the companies you want to follow
- Get ⚡ alerts when your saved companies change
- Every product they ship, cross-linked to 900+ AI tool profiles
- Today’s top AI Stories — the day’s most important AI news, free

Your AI Hub — sample view. Click to enlarge.