27 Sep 2026
Why Neon Invested

5 Min Read

Why Neon Invested in Hopper

Why Neon Invested in Hopper

“Compute is the scarcest resource today and our goal is pretty simple: make every bit of it go further,” Pavan Katta, Co-Founder at Hopper.

Most companies treat compute as a line item. Hopper treats it as an engineering problem, going below the model and the framework, down to the kernels that determine how much work a chip actually does per second. And they chose the workload where that effort pays off fastest: voice.

Voice Is the Hardest Inference Problem

Every other AI product can hide latency. A chat interface streams tokens and the user reads along. A coding agent takes thirty seconds and nobody minds.

Voice has no such cover. Human conversation has a turn-taking rhythm measured in a couple of hundred milliseconds, and the moment you exceed it, the illusion collapses. The user talks over the agent. The agent talks over the user. Everyone hangs up.

And in most production stacks, the budget has to cover three models, not one. Speech-to-text, then a language model, then text-to-speech, each handing off to the next, each adding milliseconds that cannot be recovered downstream.

So voice is where inference efficiency stops being a cost-line question and becomes a product-quality question.

The Performance Is Already There

Hopper’s engineering blog is the most convincing diligence document we read, and it is public.

Take their work on the LLM client. The standard OpenAI client defaults to HTTP/1.1 with a five-second keepalive, sensible for chat. But a voice turn takes longer than that: the agent speaks, the user replies, transcription runs. By the time the next call fires, the connection has been reaped, and you pay a fresh TCP and TLS handshake. From Europe to US-West, that is roughly 250 milliseconds of pure overhead, every turn.

They changed five settings. Median time-to-first-token fell from 510 milliseconds to 197 with conversational gaps, and from 600 to 319 milliseconds under twenty concurrent sessions.

No new model, no new hardware. Just a team that knows the voice request path well enough to know where to look.

This is the pattern that convinced us. The gains in voice inference today do not require a research breakthrough. They require obsessive engineering applied to a problem the large inference providers treat as one workload among many.

Compute Is the Scarcest Resource

Everyone building AI products today is queuing for the same scarce Nvidia chips at the same inflated prices. Hopper went and made a cheaper alternative work instead. They spent weeks rewriting the low-level code that AMD’s hardware runs, and got it to match Nvidia’s on the measure voice cares about most: how fast the first word comes back, at around 60% of the monthly rental cost.

That is not a feature. In a market where compute is rationed, being able to run on hardware your competitors cannot use economically is a structural advantage. It is also precisely what Pavan said they set out to do.

From Serving to Owning the Model

Hopper started as inference. It has since moved upstream into post-training — taking open-source speech models and tuning them on a customer’s own production calls, then serving them and improving them against live traffic.

That changes the shape of the company. Inference alone is a margin business fighting on price. Post-training on a customer’s own audio is a data loop: the longer they run on Hopper, the better their voices get, and the harder it becomes to leave. A specific accent, a specific emotional range, a specific character; things off-the-shelf TTS cannot give you at any price.

The first market pulling on this is not the one we expected. It is entertainment. Audio series, AI characters, in-game NPCs, products where the voice is not a convenience layer over a support queue but the thing the customer is actually buying.

Founder Market Fit

Pavan and Jashwanth met in high school more than a decade ago and went to IIT Madras together. When Pavan decided to start a company, he wrote, Jashwanth “was my first call.”

Pavan spent fifteen months at Salient doing exactly this, post-training and inference for voice models, at a company that raised a $60M Series A led by a16z. Before that, he was the first hire at Dashworks, building petabyte-scale indexing and retrieval before HubSpot acquired it, and before that, search infrastructure at Bing. In between, he took a year off for mechanistic interpretability research, published a paper on scaling laws and feature superposition, and was a MATS fellow under Neel Nanda.

Jashwanth wrote GPU kernels at Meta Reality Labs for over two years. Before that, he built petabyte-scale real-time streaming for Uber’s entire trip and driver dataset.

An unusually precise pairing. One has done voice post-training inside a company that sells voice agents commercially, so he knows which failures actually lose customers. The other writes the kernels.

Proud to Partner with Hopper

Voice will be how most people talk to most software. That is no longer a contrarian view.

What is still underappreciated is that it will be won on milliseconds and cost per minute, not on model quality, because the models are commoditising fast and the serving layer is not. Somebody has to make voice cheap enough and fast enough to run at consumer scale, on whatever silicon is available rather than whatever silicon is fashionable.

We believe Hopper makes voice AI viable at consumer scale. We are excited to back them.

Prasanth Garapati

Prasanth Garapati is the Vice President at Neon Fund and the designated lead for the fund's U.S. efforts. He previously founded Commut (a tech-enabled shuttle service for daily office commuters, acquired by Careem, which Uber acquired six months later) and went on to launch and scale Careem Bus across Egypt, Saudi Arabia, and Jordan. He later co-founded two B2B AI startups, flux.chat and Weka, and led U.S. and European mid-market GTM at Ema, a $140M-funded GenAI platform. He brings a decade of experience across the founder, operator, and investor seats. His investments at Neon include Hopper Inference, Mantis Grid AI, and Humpback AI.

Vector Graphic Vector Graphic

Brighten your inbox weekly with Neon’s expert insights.

Please enter a valid email id

Other posts

photo

Why Neon Invested

・ 3 min read

Why Neon Invested in Astra Security

Astra Security's pentest platform combines automated and manual penetration testing that helps businesses identify and address security vulnerabilities proactively. [...]

Read More... from Why Neon Invested in Astra Security

photo

Why Neon Invested

・ 4 min read

Why Neon Invested in Atomicwork

Atomicwork is changing the landscape of Service Management Softwares. They're building modern, easy to use, AI-first ITSM and ESM solutions. [...]

Read More... from Why Neon Invested in Atomicwork

Brighten your inbox with Neon’s insights

Brighten your inbox with Neon’s insights
Please enter a valid email id