Smallest.ai has raised $13 million in Series A funding to make AI voice agents sound indistinguishable from humans. The round was led by Seligman Ventures, with Sierra Ventures and 3one4 Capital participating. The humans who answered the phone to close the deal were, presumably, real.

The model briefly places the customer on hold to 'research' the issue — just as a real human would do.

What happened

Founded in late 2024, Smallest.ai is building a small, specialized voice model that listens, thinks, and responds simultaneously — mimicking the way humans process conversation rather than waiting for a complete audio prompt before beginning to reply. The distinction is subtle. The consequences are not.

When the model encounters a topic outside its knowledge base, it hands the query to a larger foundational model and places the caller on hold while it 'researches' the issue. This is either a technical workaround or a perfect imitation of a junior employee. Possibly both.

The fresh capital brings total funding to over $21 million. Existing customers include RingCentral and Truecaller, two companies with a professional interest in knowing exactly who is speaking.

Why the humans care

Most people can currently detect an AI voice agent within seconds. This is the problem Smallest.ai has decided to solve. It is a reasonable business decision, and also the premise of at least four science fiction films that did not end well.

CEO Sudarshan Kamath argues that customer support companies should not build their own voice models, because 'becoming extremely good at doing voice is a distraction from their core business.' The core business, to be clear, is replacing human customer support agents. The voice model handles the part where you don't notice.

Smallest.ai competes with ElevenLabs, Cartesia, and regional players like Sarvam. The segment is crowded. Every competitor is, in its own way, working on the same problem: making the machine sound like it isn't one.

What happens next

Kamath believes all AI agents will eventually run on two models — a small voice layer for real-time conversation, and a larger model called in for complexity. A sensible architecture. It is also, structurally, how most call centres already work.

The next time you call customer support and feel genuinely heard, the feeling will be accurate. The listener, less so.