Fireworks AI

Inference and fine-tuning platform for turning open-source models into specialized, production-grade AI systems.

Share

Listen to this lesson

Free preview · first 0:30
0:00 / 0:30

Unlock audio and more

Audio streaming, downloadable PDFs and certificates come with Plus and Pro.

📋About Fireworks AI

Updated August 26, 2026

Fireworks AI is an inference and fine-tuning platform that helps companies turn open-source foundation models into specialized, production-grade systems trained on their own data. Founded in 2022 by a team from Meta's PyTorch group and led by CEO Lin Qiao — formerly the Senior Director of Engineering who ran PyTorch at Meta — the company is based in Redwood City, California. Its stack centers on proprietary serving technology: the FireAttention inference kernel, speculative decoding, and the FireOptimizer adaptive engine, which together push high throughput at low latency. Fireworks serves a broad catalog of open models — including Llama, DeepSeek, Qwen, and Mixtral — through an OpenAI-compatible API, and now handles more than 40 trillion tokens a day for customers such as Cursor and Harvey. In July 2026 it raised $1.5 billion in a Series D led by Atreides Management, Index Ventures, and TCV, with backing from Nvidia.

🛠️Products & Tools (1)

Fireworks AIInference & Model Serving

Inference and fine-tuning platform for serving specialized open-source models at production scale — powered by the FireAttention kernel, speculative decoding, and the FireOptimizer adaptive engine. OpenAI-compatible API across Llama, DeepSeek, Qwen, and Mixtral.

Keep track of the companies you’re watching

Sample AI Hub dashboard showing saved tools, a content-updates alert, an AI Skill Streak, and personalized recommendations

Your AI Hub — sample view. Click to enlarge.

📰Fireworks AI in the News

Showing the only story where Fireworks AI is tagged in Top AI Stories.