Learning Objectives
- Understand why Smallest.ai builds small speech models instead of large general-purpose ones
- Identify what each of its four first-party models does and where it fits in a voice pipeline
- Evaluate its per-minute pricing against hosted competitors
- Judge which of its performance claims are independently established and which are vendor claims
What Is Smallest.ai?
Smallest.ai is a voice AI company based in Bengaluru, India, founded by Sudarshan Kamath and Akshat Mandloi, both previously engineers at Robert Bosch. In July 2026 it raised a $13 million Series A, led by Seligman Ventures, with Sierra Ventures and 3one4 Capital participating, bringing total funding past $21 million.
The company's argument is narrow and worth stating plainly: in a spoken conversation, latency matters more than raw capability. A model that is marginally smarter but answers a beat too late feels less human than a faster, simpler one. So instead of routing every turn through a large general-purpose model, Smallest.ai trains compact models that listen, reason and speak concurrently, and hands off to a larger third-party model only when a query genuinely needs it.
That design choice is the whole product. It is also the thing to test — the approach lives or dies on whether the small models are good enough for the turns they handle alone.
Key Capabilities
Lightning — text-to-speech
The flagship speech model, advertised at roughly 100 milliseconds of latency across more than 15 languages. The current generation is Lightning v3.1; the company's own documentation marks Lightning v2 and lightning-large as deprecated, so new builds should target v3.1.
Pulse — speech-to-text
Transcription across 38 languages, with speaker separation and emotion detection. Priced far below the text-to-speech side, at roughly one cent per minute.
Hydra — speech-to-speech
A native speech-to-speech model, meaning audio in and audio out without a text round trip in between. This is where the concurrent listen-think-speak design shows up most directly.
Electron — a small language model
A language model under three billion parameters, intended to handle the reasoning inside a voice turn without the cost or latency of a frontier model.
⚠️Warning
Treat the frontier comparison as a vendor claim. Smallest.ai states that Electron outperforms GPT-4.1. That claim appears on the company's marketing site rather than in a published evaluation or an independent benchmark, and a model under three billion parameters beating a frontier model across the board would be a remarkable result. Benchmark Electron on your own workload before designing around it. The latency and language-coverage figures are ordinary product specifications and are easier to verify yourself with the free credits.
Voice-agents platform
Above the models sits a configuration layer for building and deploying agents — call orchestration, a testing suite, and deployment. This is the surface most buyers actually purchase; the models are what make it cheap.
Pricing
- $10 in free credits to start
- 20 concurrent calls included
- Basic testing suite
- 1 cent per minute hosting fee
- Dedicated infrastructure with SLAs
- SSO, advanced security, compliance
- HIPAA zero-data-retention option
- On-premise deployment
The per-minute range reflects model choice rather than a plan upgrade. Speech-to-text runs around one cent per minute and text-to-speech around nine cents; if you route a turn through a third-party frontier model, that provider's cost passes through on top at cost. The company advertises rates as low as five cents per minute at volume.
For evaluation, the $10 in free credits is the relevant number — it is enough to run real traffic through the stack and measure latency on your own workload rather than trusting the advertised figure.
Smallest.ai vs. Other Voice Platforms
| Platform | Model approach | Stated latency | Languages | Commercial licensing |
|---|---|---|---|---|
| Smallest.ai | Small first-party TTS, STT, speech-to-speech and small LLM | ~100 ms text-to-speech | 15+ speech, 38 transcription | Self-serve, published per-minute rates |
| ElevenLabs | Hosted text-to-speech and voice cloning | Low-latency streaming tier | 30+ | Self-serve tiers |
| OpenAI TTS / Whisper | Hosted speech models plus a Realtime API | Realtime API for live audio | Broad | Self-serve API |
| Fish Audio | Hosted platform plus downloadable weights | Real-time text-to-speech | 30+ | Hosted tiers grant commercial use; weights are research-licensed |
| Voxtral TTS | Open-weight, self-hosted | Depends on your hardware | Varies | Open source |
The useful contrast is with Fish Audio, which also sells a hosted platform backed by its own models but publishes downloadable weights under a restrictive research license. Smallest.ai publishes no weights at all — everything is hosted API access. That removes the licensing ambiguity but also removes the self-hosting escape hatch.
Strengths
- Latency as the design goal, not a feature — the concurrent listen-think-speak architecture targets the specific failure that makes voice agents feel robotic
- Transparent per-minute pricing — published rates and $10 in free credits mean you can evaluate it without a sales conversation, which is not true of much of this category
- Full stack in one vendor — text-to-speech, transcription, speech-to-speech and the reasoning model come from one provider, so there is one bill and one latency budget
- Compliance posture aimed at regulated buyers — SOC 2 Type 2, ISO 27001, GDPR and HIPAA claims, with a zero-data-retention option on the enterprise tier
- Real telephony customers — RingCentral and Truecaller are contact-center and telephony businesses, which is a meaningful signal for a voice vendor
Limitations and Considerations
- Early-stage company — a $13 million Series A is small next to the funding behind ElevenLabs or OpenAI, and voice infrastructure is a category where vendors get acquired or repriced
- Frontier claims are self-published — the Electron-versus-GPT-4.1 comparison has no independent evaluation behind it
- No downloadable weights — hosted API only, so there is no self-hosting path and no way to pin a model version outside the vendor's lifecycle
- Deprecation cadence is real — Lightning v2 and lightning-large are already marked deprecated, so plan for migrations
- Third-party model costs pass through — the advertised per-minute rate covers the speech layer; routing turns to a frontier model adds that provider's cost on top
Company Details
Smallest.ai is a private company headquartered in Bengaluru, India. Total funding is above $21 million following the July 2026 Series A, led by Seligman Ventures. No post-money valuation was disclosed. Chief executive Sudarshan Kamath has said the aim is a voice model a caller cannot distinguish from a person.
Key Takeaways
- Smallest.ai sells small, fast speech models rather than frontier capability — the bet is that latency is what makes voice agents feel human
- Four first-party models cover the pipeline: Lightning for speech, Pulse for transcription, Hydra for speech-to-speech and Electron for in-turn reasoning
- Self-serve pricing starts at $10 in free credits with per-minute rates from about nine cents, so it is genuinely evaluable without a sales call
- The claim that Electron beats GPT-4.1 is the company's own and is not independently verified — benchmark it yourself before designing around it
- There are no downloadable weights, so unlike Fish Audio or Voxtral there is no self-hosting fallback if pricing or availability changes



