A voice-AI company building ultra-low-latency, human-sounding speech — the Lightning and Waves text-to-speech models, the Pulse speech-to-text stack, and the Atoms real-time voice-agent platform, tuned for sub-100ms conversational voice.
Last verified · 2026-07-31 · by Moe Ameen
Smallest.ai is a voice-AI company founded in late 2024 and led by founder-CEO Sudarshan Kamath. Its focus is narrow on purpose: not a general-purpose model, but voice specifically — the accents, languages, prosody, and noisy-room robustness that make synthetic speech sound like a person rather than a hold-music robot. On July 30, 2026 it announced a $13M Series A led by Seligman Ventures, with Sierra Ventures and 3one4 Capital participating, bringing total funding to over $21M.
The product line is built around speed. Lightning is its text-to-speech model, positioned as one of the fastest available, with response times measured in the low tens-to-hundreds of milliseconds; Waves is the TTS platform layer around it, offering voice cloning from a few seconds of reference audio and support for 30+ languages. Pulse is the speech-to-text side — a streaming STT stack that the company says handles 38 languages with millisecond-level latency, plus speaker diarization, emotion detection, code-switching, noise reduction, and PII/PCI redaction. Atoms is the real-time voice-agent platform that plugs those models into business phone systems for support, lead qualification, and outbound calls.
Alongside the raise, Smallest.ai debuted Voice 4.0, an architecture it frames as "asynchronous" — the system listens, reasons, and speaks in parallel the way a human does mid-conversation instead of waiting for a full turn to finish, and its Hydra speech-to-speech model is the first expression of it. Named customers span the voice and enterprise space, including RingCentral and Truecaller. Honest scope note: Smallest.ai is aimed at developers and enterprises building live voice agents and phone experiences, not at content creators. But its TTS models — Lightning and Waves — are directly useful for narration and voiceover, which is where a creator's workflow intersects it.
Smallest.ai's Lightning and Waves models solve one hard problem cleanly: turning a script into voiceover that sounds human, fast and cheap enough to do at volume. What they don't do is write that script in your brand voice, generate the video and posts around it, or put any of it in front of an audience. That's the seam Kompozy fills, and the two split cleanly at the script. Kompozy generates the on-brand script — held to your Persona Brief and banned-word filters — that you feed into Waves to voice, so the words going into Smallest.ai are already audience-fit; Waves returns the standalone narration in a cloned or multilingual voice.
From that same script, Kompozy produces everything the audio can't become on its own and publishes it. It generates Clipped and Persona Shorts for the feeds — Kompozy's talking-head video renders through HeyGen's own native voice, so it produces a face-on-camera clip without needing your import — plus brand-exact Carousel Posts via HyperFrames, Quote Graphics, Photo Posts, a Blog Article, an Email Newsletter, and Text Posts, every piece governed by your Persona Brief so the written voice matches the spoken one. Autopilot and a per-post review pipeline reframe each output to 9:16, 1:1, and 16:9 and publish across nine destinations — the eight primary social platforms plus blog and email. Smallest.ai makes the voice; Kompozy writes what it reads and ships everything around it.
Smallest.ai is a voice-AI company building ultra-low-latency, human-sounding speech models. Its lineup includes the Lightning and Waves text-to-speech models, the Pulse speech-to-text stack, and Atoms, a real-time voice-agent platform for phone support and outbound calls. On July 30, 2026 it raised a $13M Series A, taking total funding past $21M.
Not primarily. It is aimed at developers and enterprises building live voice agents and phone experiences. But its TTS models, Lightning and Waves, produce fast, realistic voiceover, which is directly useful for narration — that is the part of it a content creator would actually use.
The company positions Lightning as one of the fastest text-to-speech models, with latency measured in milliseconds, and says its Pulse speech-to-text handles 38 languages with millisecond-level latency. Waves supports voice cloning from a few seconds of audio across 30+ languages. Treat exact numbers as vendor-stated.
They split at the script. Kompozy generates the on-brand script you voice in Smallest.ai, then turns that same script into finished content — Clipped or Persona Shorts (its talking-head video uses HeyGen's native voice), a carousel, quote graphics, a blog, and a newsletter — and publishes across nine destinations: eight social platforms plus blog and email.
Smallest.ai competes with voice-AI providers including ElevenLabs and Cartesia, and regional players like Sarvam. For creators, the more relevant comparison is other TTS tools you can feed into a content pipeline — the voice model is only one step; a tool like Kompozy is what writes the on-brand script and generates and publishes the video and posts around it.