Cartesia logo
AIvoice-aigenerative-aistate-space-models

CARTESIA

Netfigo Verdict
on Cartesia

Cartesia is what happens when the people who tried to dethrone the transformer point their weird new AI architecture at voice. Co-founder Albert Gu helped invent Mamba, a state space model that runs faster and cheaper than the tech behind most chatbots. In 2023 he and a crew from the Stanford AI Lab spun out to build Sonic, a voice model fast enough to talk back in real time. Kleiner Perkins led a $64 million Series A in 2025. The bet is contrarian and technical: that the future of AI voice runs on a different engine than everyone else is using.

Founded

2023

HQ

San Francisco, USA

Total Raised

$191 million

Founder

Karan Goel, Albert Gu, Arjun Desai, Brandon Yang, Chris Ré

Status

Private

THE ORIGIN STORY

The story starts in a Stanford lab, not a garage. Albert Gu spent his PhD years working on state space models, an alternative to the transformer, the architecture behind almost every big AI model today.

In 2023 he co-authored Mamba, a design that handles long streams of data faster and with less memory. That is a huge deal for voice, where the audio never stops and every millisecond of delay is noticeable.

Gu teamed up with Karan Goel, Arjun Desai, Brandon Yang, and Stanford professor Chris Ré, a MacArthur genius grant winner. They spun their research out of the Stanford AI Lab and founded Cartesia.

The mission was blunt. Build real-time AI on a better engine, starting with voice.

WHAT THEY ACTUALLY DO

Cartesia sells voice as an API. Developers plug it in and pay for what they use.

If you are building an AI phone agent, a game character that talks, or an app that reads text aloud in a natural voice, you can use Cartesia instead of training your own models. The whole pitch rests on two words: fast and cheap.

Because Cartesia's voice model Sonic runs on state space models instead of the standard transformer, it can generate speech with very low delay. Low delay is the difference between a conversation that feels human and one that feels like a bad phone menu.

THE PRODUCTS

The flagship is Sonic, a generative voice model built for real-time speech. It can clone voices and produce natural-sounding audio with very low latency, which is the whole point for live conversations.

Cartesia has kept pushing the model forward, later releasing Sonic-3 alongside a $100 million raise. Beyond the voice model, the company is building a broader platform for real-time AI, all resting on the state space model research its founders are known for.

The through-line is speed. Everything is designed to respond instantly.

HOW THEY GREW

Cartesia's edge is not marketing. It is the underlying tech.

Most voice startups build on the same transformer architecture, so they compete on data and tuning. Cartesia went a different route with state space models, the design Gu helped invent, which lets Sonic respond faster and cost less to run at scale.

That is a real technical wedge in a crowded market. The company also leaned into being the credible research shop.

When your co-founders literally wrote the papers behind the architecture, developers and investors take the speed claims seriously. It landed serious money fast, going from a $27 million seed to a $64 million Series A within months.

THE HARD PART

Voice AI is brutally competitive. ElevenLabs, OpenAI, Deepgram, Google, and a pile of funded startups are all racing for the same developers.

Cartesia's bet on state space models is smart, but it is also unproven at massive scale compared to the transformer, which the entire industry has optimized for years. If transformer-based rivals close the speed and cost gap, Cartesia's main advantage shrinks.

The company also has to prove that being the fastest voice model translates into real revenue, not just impressive demos. Being technically brilliant and being a durable business are two different things.

MONEY TRAIL

Seed

2024 · Led by Undisclosed

$27M raised

Series A

2025 · Led by Kleiner Perkins

$64M raised

WHO BACKED THEM

Cartesia raised a $27 million seed in December 2024, then a $64 million Series A led by Kleiner Perkins in March 2025, with Index Ventures, Lightspeed, and others joining. It later added a $100 million round tied to the launch of Sonic-3, bringing total funding to roughly $191 million.

The backing is a bet on the founders as much as the product. When the people building your voice model are the same researchers who invented the architecture underneath it, investors treat the technical claims as credible rather than hype.