I'm an ambitious product leader. Right now I lead the voice agents charter at Sarvam. I like building products people actually enjoy using, and then getting a lot more people to use them. Growth is the part I find most fun.
Before Sarvam I spent three years at Fi Money running wealth: US Stocks, mutual funds and the net worth tracker. Before that, startup consulting at BCG, working on growth and go-to-market for early-stage companies.
I quiz, and have since school. It's left me with very varied interests and a lot of random knowledge across a lot of domains.
I read a lot too: business, geopolitics, biographies, sci-fi and fantasy. Fan of all things nerdy, Harry Potter, LOTR, Star Wars, the lot.
And robotics, always. From Lego robots in school to computer vision in college.
02 · Work
Where I've been
Sarvam AI
Product lead, Voice Agents2025 to present
#1Voice AI in Indiathe largest deployment in the country
22xGrowthin daily voice minutes since I took the charter
>2.5MVoice calls a dayreal conversations, not API calls
3xCheaper per minutecost of a voice minute, down 3x in the same period
Launch · Feb 2026 · India AI Impact Summit
Sarvam 30B, launched with a phone call
Ran the 30B release on the main stage at the Summit
Demoed it live by calling from a basic keypad phone. No app, no internet
The voice on the other end was Vikram, our assistant for feature phones. Named after Vikram Sarabhai
Bulbul v3 shipped alongside, with 35 voices. Helped drive the virality around it. How we ran evals
ProductFig. 02Home. What should your voice agent do?Fig. 03The canvas. Instructions, variables, tools, tests.Fig. 04Workflows. An API call, a decision, then a phone call, drawn as a flow.Fig. 05Phone numbers. Rent one, or bring your own telephony.Fig. 06Analytics. Goals, turns, latency, language split. Numbers blurred on purpose.
22x growth in daily voice minutes since I took the charter
Largest voice AI deployment in India
More than 2.5M voice calls a day, in 12 languages
Cut the cost of a voice minute by 3x along the way. Happy to talk it through on a call. Email me
Daily voice minutes, month by month. Real shape, axes deliberately left off. The gold dot is the day we crossed a million minutes.
While you've been on this page
0
voice calls answered by agents I work on. Around 55 every second of the working day.
Fi Money
Fi set out to make money fun and useful for people in their twenties: a neobank with a personality, built on top of a real bank. It raised about $240 million from Peak XV, Ribbit, B Capital, Temasek and Alpha Wave to do it. I ran wealth, the part that had to make people richer, not just organised.
In 2026 Fi wound down its consumer business. Regulation favours the banks that already exist, trust is slow to earn, and it turned out to be hard to make money on anything but lending.
Associate Director, Product · Wealth2024 to 2025Senior PM · Wealth2022 to 2024
1 in 5Daily users, on net worthabout 20k people a day, out of roughly 100k
₹3,000 Cr+Assets trackedon the net worth tracker
1stPersonal finance MCPin India. Your whole net worth in Claude or ChatGPT, with a single click
50k+US brokerage accountsand $5M+ in forex, in the first year
Wealth · 2024 to 2025
Net worth, rebuilt around a morning ritual
"I wake up in the morning, and the first thing I do is check my net worth."A power user, in a call. It wasn't about the number. It was about control.
Took mutual fund connection from ~2% to ~16%. Cut acquisition cost 8x
Daily portfolio tracker. How you did, every morning. Retention roughly doubled on the back of it
Fig. 08
Home, a money secret, and the insights page. The EPF one found interest people didn't know they were missing.
Launch · Jul 2025
Fi MCP. Talk to your entire net worth
An industry first: one MCP over all of it. Mutual funds, Indian stocks, bank accounts, credit score, loans, EPF
Works with Claude, ChatGPT, Gemini, Cursor. Or tap "Talk to AI" in the app
Secure, plug and play, answers to real money questions in under 3 minutes
"Which loan EMI should I repay first?""Can I take a career break?""Is my portfolio diversified?""How do I get my credit score to 800?"
Product · 2022 to 2024
Invest in US Stocks from India
The hard part was the remittance. We made sending money to the US work like a UPI payment: a few taps, and the dollars are in your account within a day
KYC, the remittance and the US broker, all inside one app
Fractional shares from ₹100. Zero brokerage, zero account fees
Sign up to first trade in about 4 minutes
50k+ brokerage accounts and $5M+ moved in the first year
Fig. 10
Fig. 11
Adding dollars, and what people bought with them.
Growth · Mutual funds and deposits
Rules that invest for you
Auto-investment rules: IPL rules, Daily SIP, Invest the Change, AutoPay. ₹50 crore+ of new money
SIPs from ₹10. Three times as many first-time mutual fund investors
Cut pre-closure churn on deposits by 10%, about ₹10 crore of AUM a year kept in
Fig. 12
BCG
Startup consulting in BCG's Growth Tech practice. Ten months, four clients, top-bracket ratings every cycle. I left to build and operate on the ground rather than advise from 30,000 feet.
Confidentiality keeps the details off this page. The shape of it:
Senior Associate, Growth Tech2021 to 2022
A 10x growth plan for a large gig-services company's painting business. Our one-day pilot did 1,000+ jobs in two months
Growth strategy for a D2C personal care brand
Commercial due diligence on a $400M acquisition, for a D2C eyewear brand
Product roadmap and investor deck for an edtech company
School
MBA
IIM Ahmedabad
Institute Scholar (top 2.5%). Not the plan, went well anyway.
B.Tech Mechanical
IIT BHU
Spent most of it on robots, dynamics and machine learning.
IIM Ahmedabad · IIT BHU2015 to 2021
03 · ProjectsNights and weekends
Things I did outside of work
Voxforge.ai
Built Dec 2025 · Live
Mission control for audio production. Nine TTS providers, one interface, a casting director that auditions voices for every character in your script.
The landing page. Mission control for audio production.The casting director auditioning three voices for a narrator.The workbench. Takes per line, promote the winner.The timeline. Dialogue, effects and music, in the browser.
What I learned
People loved take management and promoting the best take. It's a feature, not a product.
Being a bit better than ElevenLabs wasn't enough. People generally just preferred ElevenLabs.
The companies I spoke to didn't want to be limited to audio. Most of the asks were for the same thing, for video.
Does
Casting director · script parser · takes · timeline · team review
A college project. A small snake robot, twelve 3D-printed modules on Dynamixel servos, built to move over sand where wheels get stuck. We programmed three gaits into it: linear propagation, rolling and sidewinding.
Moving, on the lab floor.Tethered, mid-gait.The CAD it started as.
Built
12 modules, Dynamixel AX-12A servos, 3D printed
Gaits
Linear propagation · rolling · sidewinding
Pitstop
Side venture · 2023
Firms make employees promise not to trade on inside information, then have no way to check. Pitstop pulls their real trades from all 300+ brokers over the RBI's Account Aggregator rails, with consent, and screens them against the firm's policy.
The solution. One consent, every broker.The problem we were selling against.Where the data sits, and who can see it.
What I learned
Enforcement in India is extremely weak. Firms will buy a tool that proves they tried, not one that actually catches people.
Compliance software sells to a team with no budget and no upside for saying yes.
Rails
RBI Account Aggregator · SEBI RIA and PFRDA RA licences pending
A disgraced physicist. A quantum accident. And a garden snail that will follow him to the end of time. A saga that spans five thousand years. Out of the Reddit immortal-snail thought experiment.
THE SNAILAditya Dhavala
Sometimes the most elaborate prisons are the ones we meticulously build for ourselves, believing them to be pathways to salvation.Chapter 16
You cannot cheat the universe of its balance.Chapter 16
Immortality had taught him the value of something finite.Chapter 5
Size
16 chapters · 405 pages · ~57,000 words
Owes
Watchmen, The Three-Body Problem, one Reddit thread
A short-form audio app I co-founded in 2021, in the year everyone thought audio was next. It wound down when the engineers left for other things. 2021
CObRaSO
A crawler module that bends in any direction and keeps crawling while bent. I worked on it at the Robotics Research Center at IIIT Hyderabad, and I am an author on the original paper. 2017
All of these I've read and would hand to you. Hover a spine.
+150
The snail book. Click it.
May 2026Voice
Speech to speech vs cascaded pipelines, from production
Speech to speech models feel magical for about a week. Then you try to ship one. Cost, control, attribution, the improvement loop, and the harness you end up building anyway.
Speech to speech models feel magical the first time you use one. The latency is gone, the prosody is right, it laughs in the right places. If you have spent years assembling voice agents out of three separate models, the first good end to end demo feels like the whole category just got obsoleted.
Then you try to put one into production, and it gets painful in a very specific set of ways.
Two caveats before I start. This is my own view and not my employer's, and nothing here represents anyone's position but mine. And I may be eating these words inside a year, so I have tried to keep the structural arguments separate from the ones that are only about where the models happen to be this month.
The short version
Cost. Voice agents are commoditising, so the low cost producer wins. Speech to speech is several times more expensive per minute.
Control. Past a few thousand tokens of context, instruction following starts to slide. A single tool call returning a real payload can put you over the line.
Attribution. When a call goes wrong, a cascade lets you say which of the three models broke. With speech to speech you have the recording and the outcome, and no way to get between them.
The improvement loop. Text traces can be read, judged and corrected by typing. Audio traces cannot, so the flywheel turns slower.
The harness. Getting a speech to speech model to behave reliably takes a large amount of custom work, and that work is not portable. It is shaped around one model's limits, so it carries over to nothing else.
The two things
A cascaded pipeline is three models in a line. Speech to text, then a language model, then text to speech. Each hands off to the next, and text sits between every stage.
A speech to speech model takes audio in and gives audio out. There is no text in the middle.
Almost everything below follows from that. Text in the middle is a seam, and a seam is where you put a cache, a log, a test, a fix, and a different model.
Cost
The voice agents market is commoditising quickly, and I have written separately about why. The short version is that there are no real network effects, feature differentiation lasts about a quarter, and the economies of scale accrue to whoever owns the model rather than whoever owns the application. In a market like that the low cost producer wins.
So cost is not a line item here. It is the strategy. Three numbers:
The market is selling voice agents at around two rupees a minute.
A team orchestrating a cascade sits well under a rupee, and that gap is where the margin lives.
Update, June 2026. Google's release moved this a little. Gemini's native audio Live model publishes half a cent a minute for audio in and under two cents for audio out, and the mini tier of OpenAI's realtime model is in the same territory. On paper that is competitive with a cascade, which is not where the cheap end of this market sat when I first wrote this.
Two things to say about that.
The first is what those numbers count. They price a single pass of audio. A real conversation re-sends the accumulated context on every turn, and Google is explicit that the Live API re-counts everything, with a Google engineer stating on their own forum that all tokens are counted at every turn, including tokens from the previous context. So you buy the same audio over and over, and the effective rate sits above the sticker and grows with the length of the call. How far above depends on how much history you keep and whether you keep it as audio or as text, and that is a decision a cascade lets you make and a realtime API mostly does not. The published per-minute figure is a floor you would only ever hit on a conversation one turn long.
The second is that I think the cheap tiers are priced to take the market rather than to cover what they cost to serve. That is a guess and I cannot prove it from public filings, but the arithmetic points at it. Audio is generated at a frame rate, so a minute of speech is a few thousand output tokens where the same minute of text is a couple of hundred. OpenAI's own documentation puts the multiple at roughly ten times more tokens for the same sentence in audio than in text. Ten times the decode steps to say the same thing. There is no architecture in which that is cheaper to serve than running a small text model and then a text to speech pass.
The tell is at the top of the range, where nobody is buying market share. OpenAI launched gpt-realtime in August 2025 at sixty-four dollars per million audio output tokens, and it is still sixty-four through two revisions of the model since. Google has held twelve dollars per million for native audio output across every generation it has shipped. And nobody has cut the price of a native audio model already in the field since OpenAI's sixty percent cut in December 2024. New models have arrived cheaper, which is a different thing. What nobody does is drop the price of the model you are already paying for. A price that will not move at the top of the range, at two vendors, in a market where everyone is fighting for share, usually means you have reached the real cost of decoding. The mini tiers sitting well below that line look like customer acquisition.
If that is right, the interesting number is not today's list price. It is what the price has to be eventually, and that is set by tokens per second of speech, which favours the cascade permanently.
Control
Past about five thousand tokens of context, a speech to speech model starts to misbehave. Instruction following slides, it begins ignoring parts of the prompt, and behaviour stops being repeatable.
Tool calling itself works fine. What breaks it is a tool that returns something real, a customer record or an order history, because that arrives all at once and puts you over the line inside a single turn.
That threshold is from our own testing rather than from anyone's documentation. The published numbers explain why it lands about there.
But audio runs about ten times the tokens of the same content in text, by OpenAI's own reckoning. So 32,000 audio tokens is somewhere near three thousand text equivalent tokens of dialogue.
Then take out system instructions, capped at 16,384 tokens on that API.
And the roughly four thousand reserved for output.
And your tool schemas.
A practical ceiling of a few thousand tokens is what that arithmetic predicts, and it degrades well before the window fills. OpenAI says so in their own cookbook: in certain use cases you may notice performance degrade as you stuff more tokens into the context window. Their suggested fix is to start summarising somewhere between twenty and thirty-two thousand tokens. Developers on their forum report around 17,800 tokens working and 31,300 returning a 504, along with a roughly ten percent tool calling error rate, which one of them points out is not good enough when the agent is taking real orders for real money. AWS, writing about their own model, notes that sub-agents returning large raw payloads are not ideal for voice and recommends splitting a conversation into separate sessions per phase.
Meanwhile a text model handles 32,000 tokens of actual text without anyone writing a blog post about it, and for offline work reasoning models take millions of input tokens.
What the public benchmarks show
I am saying this from our own testing and internal evals rather than from anyone's leaderboard. The public benchmarks that exist point the same way, for whatever they are worth.
Reasoning is no longer the problem. Big Bench Audio is saturated. The top audio models sit at 96 to 99 percent, including gpt-realtime-2 and Gemini 3.1 Flash Live at their highest settings. Anyone arguing that audio models cannot think is a year out of date.
Instruction following and multi-turn is where it shows. Scale's audio MultiChallenge runs human speech with disfluencies and interruptions over three to eight turns, and a task only passes if every rubric is satisfied. gpt-realtime-2 scores 48 at its highest reasoning setting. Gemini 3.1 Flash Live scores 36 with thinking on and 27 with it off. The text version of the same benchmark tops out around 75.
One thing worth noticing in those numbers. The gap between thinking and non-thinking on the same model is larger than any gap between vendors, which suggests the constraint is how much the model is allowed to deliberate before it has to start talking. In a voice call the answer is not much.
Instruction following is not a nice to have in a voice agent. It is the product. It is the difference between an agent that follows your policy and one that improvises on a call you are answerable for.
The other half of control is that the voice itself is variable in production in a way a dedicated text to speech model is not. It does not always sound right, and when it does not you have no knob.
Attribution
When a call goes wrong someone asks what happened, and depending on the customer you may have to answer in writing.
In a cascade that takes ten minutes, because you can tell which of three things broke:
The speech to text misheard the account number.
The language model lost the thread.
The text to speech mangled a name or read a number wrong.
Three failures, three owners, three fixes, and you know which one you are looking at. With speech to speech you have an audio file and a bad outcome. You can listen to it, which does not scale, and you can hope the next model release is better, which is not an answer.
The least conflicted evidence is not an opinion piece, it is LiveKit's own documentation noting as a plain capability limit that realtime models do not provide interim transcription results, and that user transcriptions can be considerably delayed and often arrive after the agent has already responded. You cannot get a trustworthy trace out of the model, because the trace was never the representation.
The improvement loop
This is the same argument as attribution, one time scale up.
Debugging is working out why one call went wrong. Improving is working out why a thousand went wrong. Both need the same thing, a record of what happened inside the call.
A cascade produces that record for free. Every call leaves a transcript, and a transcript can be read in seconds, searched, grouped by failure type, and turned into a test case. When you find a bad one you can write down what the agent should have said instead, and now you have an example to fix a prompt against or to train on.
Speech to speech leaves you the audio. You can transcribe it afterwards, but that transcript is a reconstruction rather than the thing the model actually worked from, and it still does not tell you which part of the model went wrong. Review is also much slower, because reading a thousand transcripts is an afternoon and listening to a thousand calls is a week.
Underneath this is a data problem. There is far less conversational speech in the world than there is text, and the gap is not close, which is why speech models improve more slowly than text models on the same amount of compute. Cuervo and Marxer put the difference at up to three orders of magnitude.
The fair counterargument is that almost every speech to speech model is a text model with an audio front end attached, so it inherits some of the text side progress for free. That is true. What it does not inherit is the audio specific part, and the field's own answer to the audio data deficit is to lean harder on text, which is an odd place for the end to end thesis to stand.
The harness
Because the usable context is small, you cannot just point a speech to speech model at your problem. You have to build a harness around it, and the harness is where the trouble is.
I know a team doing this properly. Sub-agents, and inside each sub-agent multiple state prompts capped at around two thousand tokens each, orchestrated with a mix of deterministic rules and tool calls. That works. It is also a bespoke context engineering layer built to one model's failure profile.
The obvious objection is that the frameworks solve this, and they do not. LiveKit does abstract eight different realtime providers behind one interface, so swapping vendors at the session layer is a plugin change. But that is not the layer the lock-in lives at. Nothing in the framework holds:
your two thousand token state machine,
your sub-agent boundaries,
or the summarisation thresholds you tuned by watching one specific model degrade.
All of that sits above the framework and all of it is calibrated to a model. Change the model and the calibration is wrong, which means you re-derive it by hand.
The orchestration you end up writing is thick. Not a wrapper around an API, a whole layer: context budgets, when to summarise, when to reset, which state the conversation is in and what the model is allowed to see in each one. That layer is where all your reliability lives, and you derived every threshold in it by watching one model fail.
The pace on the other side
Now put that next to what happened to the text agent stack in eighteen months.
Tool calling standardised.
MCP went from an announcement in late 2024 to adoption across every major cloud inside a year, with the eval infrastructure versioning alongside it.
A 9B open weight model now scores around 0.66 on the Berkeley function calling leaderboard, roughly where the best frontier model sat in mid 2025.
Gemma 4 arrived in March 2026, and its 31B model is a genuinely good brain for a voice agent. Open weights, on your own hardware, swappable in an afternoon.
Every one of those is a drop-in swap for a cascade. Not one of them is available to a speech to speech model. That is the asymmetry that matters most over a two year horizon: the cascade inherits everything that happens to text models by changing one line of config, and the end to end model inherits only what its own lab ships.
It is worth noticing where the most valuable company in this market landed. ElevenLabs says it on their own blog: at ElevenLabs, we use an advanced cascade-based architecture, with specialised components for recognition, reasoning and generation, co-located in one stack so there are no network hops between them. They built their own turn-taking model rather than buying turn-taking end to end, and they went as far as hosting language models inside their own path to kill the hop. When the company whose entire identity is voice does not go end to end, that is information.
Where speech to speech genuinely wins
On the things it is good at it is not close.
The conversation layer
Turn-taking feels human. Knowing when someone has finished speaking is a prosodic judgement, not a textual one, and a cascade has usually discarded the prosody by the time it needs to decide.
Barge-in is better. Interrupting it feels like interrupting a person.
Backchannels are natural. The small "haan" or "mm" that tells someone you are still listening, rather than a scripted filler.
How you said it survives. A cascade throws away tone, hesitation and emphasis at the speech to text step, and everything downstream is working from a flat string. An end to end model still has all of it when it decides what to say.
The latency floor is genuinely lower. Two handoffs removed.
And an entire class of product
Everything I have argued against speech to speech is an argument about business calls at scale. Read the objections back and they are all the same shape. Cost per minute multiplied by millions of minutes. An audit trail somebody will ask for in writing. A policy the agent must not improvise around. A tool returning a real customer record.
A personal assistant inverts every one of them.
Cost per minute barely matters when it is one person having a few conversations a day rather than a contact centre running a million.
Nobody is going to ask you to attribute a failure in a conversation between someone and their own assistant.
The tools are light. A calendar, a timer, some music, a search. Not an order history, and not a fleet of sub-agents.
And what decides whether someone keeps using it is not accuracy. It is whether talking to it feels like talking to someone.
That last one is where the whole argument flips over. On a collections call a warm interruption is a nice touch. In a Jarvis it is the entire product. The reason to build an assistant you talk to all day instead of typing at is that it hears how you said something and answers accordingly, and a cascade has thrown that away before it starts thinking.
So I do not think these two architectures are competing. They are for different products. Transactional voice at scale wants a pipeline you can price, inspect and improve. An assistant you live with wants a model that never leaves the audio. A lot of the disagreement I see is two people arguing about different things.
Where I actually land
For a transactional voice agent, which is what I build, the answer is cascade for the content and something end to end for the conversation layer.
Let the pipeline decide what to say, because you need to read it, cache it, evaluate it, attribute it and defend it. Let a model that never discards the audio decide when to say it, when to stop, and when to make the small noise that means keep going.
Turn-taking is the one part of this problem where text was never the right representation, and it is also the one part where nobody needs an auditable artefact. That is roughly where the better platforms have converged, and it is not a compromise. It is putting each decision where its information actually lives.
What would change my mind
Cost. If absolute costs fall far enough that a three times difference stops mattering in rupees, the argument dies regardless of the ratio. That is not happening yet, given ten months of flat flagship pricing.
Control. A native audio model that holds instruction following through a long call with real tool payloads, measured on the multi-turn benchmarks rather than the reasoning ones. The reasoning gap closed faster than anyone predicted, so this one deserves humility.
Attribution. Any inspectable intermediate, even a lossy one, that lets me pin a bad call to a stage. I do not think the labs building these models consider that their problem, which is the real reason I am not expecting it.
Written May 2026, with a June update in the cost section. List prices move fast in this space, so treat anything quoted here as of those dates.
Jul 2026Opinion
Voice agents are commoditising
The consensus says prices settle at two rupees a minute. I think it is one, inside eighteen months, and the winners will be whoever gets to scale first.
Apurv Agrawal, who runs SquadStack, said recently that voice pricing in India is going to stabilise at two rupees a minute.
I do not buy it. My sense is it goes to a rupee inside eighteen months.
It was three rupees a minute last year. It is two now. I am fairly willing to bet on one.
This is my own view and not my employer's.
Why it keeps falling
Nothing to do with this year's models. It is the shape of the market. Run the five forces over voice agents and four of them come back strong.
Buyers
Almost everyone buying a voice agent is buying it for a cost centre. Contact centre optimisation is the use case by a mile, and the person in the room has been told to spend less than they spent last year.
You do create value for them. You just cannot get paid for it.
You price per minute. So does everybody else in the bake-off. So the minute becomes the thing they negotiate, and whatever value you created above the price of a minute you have handed over for free. Does not matter how good the agent is. The invoice says minutes.
Outcome pricing is the obvious escape and it is hard, because to charge for an outcome you need to own enough of the workflow to prove one happened. I will come back to that.
People will tell you the buyers are not only contact centres any more. Sales teams are using voice AI, marketing is using it, the business side of the house is waking up to it. Fine, and it is still tiny next to contact centre spend.
Suppliers
If you are a wrapper, have a look at who you buy from. Two or three text to speech models worth using. A couple of speech recognition models. That is your entire supply base.
The worry everyone has is that they put your prices up. They will not. Input prices keep falling and they will keep falling.
The real problem is that your supplier can build your product. They already own the expensive bit, so bolting orchestration on top of it is not a serious engineering project for them. When they do it they can price it somewhere you cannot go, because they are not paying anyone's margin but their own, and they get better cache hit rates because they own the whole path. Then they quietly move the volume onto themselves.
And this is not a hypothetical. ElevenLabs sells an agents product. Cartesia sells one. Deepgram and AssemblyAI both ship one. Every serious model company in this market has already built the layer above itself, which tells you roughly what they think of the difficulty.
New entrants
Anyone with a bit of engineering ability can wrap three models and call it a platform. Honestly, orchestrating those three models is a trivial product.
Getting it to hold up at scale is not trivial at all. That is where the real work is. But none of that is visible in a demo, and the new entrant does not need to have solved it to walk into your deal and quote a number.
Substitutes
This one is weak.
The substitute is a person. Move the contact centre to a tier 2 city and a human minute costs less than people think. Or the big contact centres get good at AI assisted service, with orchestration behind them and humans on the calls that matter, which some of them are already doing.
I do not think it caps the price. It does make it hard to tell a story where prices go back up.
Rivalry
Here is the force that does the work, and it is savage because nobody can differentiate.
Sell a fully customisable agent by the minute and the difference between a lead qualification agent and a debt collections agent is a prompt. Same stack, same telephony, same models, same price. Try building a differentiation story on top of a prompt. It does not survive the first conversation with a buyer who has written a few prompts himself.
Features will not save you. Every problem statement in this space is already known. Interruption handling, language switching, voicemail detection, silence nudges, warm transfer. There is no secret list going around. Everybody is building the same one, and you end up with a market of products that demo identically.
Nothing to differentiate on, and anyone can walk in. All that is left is price.
So it is a scale game
None of this means the market stays crowded, by the way. I actually think it consolidates into a few players, and most of the companies in it today will not be around in three years.
Once price is the only axis, the winner is whoever can serve a minute for less than everyone else can. And your cost per minute drops as your volume goes up. Volume gets you inference deals nobody else is being offered. Past a certain point it justifies hosting your own models, which is cheaper again. It also spreads the boring engineering that keeps a million calls a day from falling over, which is the part nobody sees and the part that actually costs money.
Then it compounds. Cheaper wins you volume, volume makes you cheaper, and everyone else is out there quoting off a cost base they do not control against someone who does.
Where the value goes
Most of it goes to the customer. That is what commoditising means. The enterprise that used to burn a hundred rupees on a call now spends two, and they keep nearly all of the difference.
Some goes to the model companies. They own the scarce thing and can come down the stack whenever they want to.
Very little goes to the wrappers.
If you are building here, that ordering is the thing to sit with. It is not about who wins. It is about which layer the money settles in, and it is not the application layer.
The way out
Not a better voice stack. Being somewhere the minute is not the unit.
The vertical players are in a completely different game. Voice sits inside a product that sells something else, as a cost line or an ROI driver, and they are not quoting per minute against four competitors on a Tuesday afternoon. The value it throws off lands on their own books instead of getting negotiated away.
Same conclusion as the outcome pricing thing earlier, from the other end. You can only charge for an outcome if you own enough of the process to prove one, and if you own that much of the process, were you ever really a voice company?
A general purpose voice stack sold by the minute into somebody else's workflow is going to keep feeling the buyer power. I cannot see a version of that business where it stops.
Is this an India story
Partly. Less than people assume.
India is more price sensitive, no question. You drop prices for volume here in a way you probably do not have to elsewhere, and willingness to pay for the same minute is lower.
But go back through the arguments above and I cannot find anything that is structurally different. Buyers are cost centres in both places. The differentiation problem is a prompt in both. The supply base is the same two or three companies in both. Barriers to entry are on the floor in both. A minute of voice audio is a minute of voice audio.
What differs is willingness to pay and how big an ROI you can point at, and that is a gap that closes, not a moat. If the price is two rupees here, there is nothing fundamental separating what an Indian buyer is asking for from what a buyer anywhere else wants. Right now the difference is mostly the ability to sell.
Going global will improve your margins for a while. I do not think it changes where any of this ends up.
Written July 2026. Prices in this market move fast, so treat the numbers as of that date.
Jun 2026Opinion
Why easy to evaluate work has no moat
Harnesses keep getting thinner. Anything you can write an eval for, the model will absorb. So the only durable alpha is in work that is expensive, undefined or tasteful to judge.
Look at what a harness looked like eighteen months ago and what it looks like now.
Eighteen months ago you were hand building the scaffolding. Prompt chains. A state machine you wrote yourself. Output parsers, because the model would not reliably give you JSON. A retry loop around a tool call. Chain of thought stuffed into the prompt because the model would not think unless you told it to, step by step.
Almost all of that is gone. The model reasons on its own now. It calls tools natively. It gives you structured output because you asked for a schema. Every one of those things was somebody's differentiated product at some point, and now it is a checkbox in an API.
The harness keeps getting thinner. I do not think that stops.
Why it keeps thinning
Whatever you build around the model becomes training data for the next model.
That is the whole mechanism. You wrote a scaffold because the model could not do something. Your scaffold produces traces of the model doing that thing correctly, over and over. Those traces are exactly what you would need to teach a model to do it without the scaffold.
Today's harnesses have sub-agents doing orchestration. Planners handing off to executors, evaluator loops, routing layers. My guess is those traces end up inside the model too, and reasonably soon. Nothing about sub-agent orchestration looks structurally different from chain of thought, which was also a technique you had to apply from outside until it was not.
So where is the alpha
If the harness thins out to nothing, what is left of a vertical product? Today I think it is three things.
Proprietary skills. You know how to do the job. The domain knowledge that is not written down anywhere, encoded as prompts and rules and workflows.
Proprietary data, or tools that reach proprietary data. Not the model's knowledge, yours. The integration into a system nobody else can see into.
A specific orchestration shape. Generator and evaluator pairs, particular decompositions, the architecture you worked out by trial and error for this one problem.
Those are real. They are producing real margin right now.
But all three have the same weakness, and it is not the one people expect.
The eval is the leak
Every one of those three can be hill climbed, as long as you can measure whether the output was good.
Write the benchmark, and an RL loop will max it out. That is the entire game now. Define the objective well enough and the optimisation finds the top of the hill without needing to know anything about your domain. Your proprietary skill becomes a number to climb. Your orchestration shape becomes something the model learns to do internally because it was rewarded for the outcome your orchestration produced.
Which gives you an uncomfortable test. If you can write a good benchmark for your own product, you have written the specification for whatever replaces it.
The corollary is the interesting half. Where the loop cannot close, the alpha stays.
RL loops get hard when:
The feedback is expensive. You cannot run a million rollouts if each one requires a physical action, a lab result, a legal filing, or a human expert reading it carefully.
The feedback is slow. If you only learn whether the answer was right in six months, there is no gradient to follow this quarter.
The objective is undefined. Not hard to hit. Genuinely unclear what hitting it means, and different experts disagree.
It comes down to taste. Someone looks at the output and knows it is wrong, and cannot tell you which term in the loss function that was.
In those places data suddenly gets more valuable, because it is the only thing that substitutes for a feedback loop you cannot run. If you cannot climb a hill, you have to have been there before.
Picking a vertical
Most people pick a vertical on size, or on how hard the work looks. Hard does not mean defensible. Plenty of genuinely difficult work has a clean right answer at the end of it, and a clean right answer is exactly what an RL loop needs, so that work goes first.
What I would ask instead is how much it costs to find out whether the job was done properly. Not whether the job is hard. Whether checking it is hard.
If checking is cheap, what you have is a head start. Worth having. Worth pricing like a head start.
This cuts against me too
Voice agents are fairly easy to evaluate.
Word error rate. Latency. Did the call complete. Did the customer get transferred. Did the lead get qualified. Almost all of it is measurable and much of it is measurable automatically, which is why I think that market commoditises, and I have written that separately. The framework does not spare my own patch, and if it did I would not trust it.
Where this argument is weak
Two places.
Hard to evaluate might just mean not yet evaluated. People said design taste could not be measured, and then A/B testing showed up and measured quite a lot of it. Evals get invented, and every time one does, a moat that felt structural turns out to have been a gap in the tooling. So the honest version of my claim is that evaluability is a rate limiter, not a wall. It buys years, not forever.
And the moat comes with a ceiling. If nobody can measure whether you are good, that includes your customers. Taste based businesses are famously hard to scale and hard to sell, because the thing you are selling cannot be demonstrated in a bake-off. So you get a defensible business that grows slowly, which is a real trade and not obviously the one you want.
Neither of those changes the test. They just mean it is a test for where value survives, not a recipe for a large company.
Written June 2026.
Nov 2025Teams
High trust, high humour teams are the actual job
A team that is not enjoying the work cannot hold pace. Weekly betting on the north star, not shipping most of the time, and why PMs and engineers should be able to do each other's jobs.
Two companies I have worked at have made Slack stickers of me. Different teams, no overlap, nobody coordinating it.
Fun is a pace requirement
A team that is not having fun while doing the work cannot move at pace for very long. Not a week or a month. It will do it for a quarter and then something goes quiet.
So a decent chunk of my job is working out where each person finds the joy in this. For some people it is the hard technical problem. For others it is watching a number move, or being the one who found the bug, or shipping something their family can use. It is rarely the same thing twice and it is almost never the thing on their job description.
The best mechanism I have found for this is betting.
Every week the team bets on where our north star metric will land. Everyone puts a number in. We run the whole thing as a ritual, with the reveal and the arguing and the winner being whoever came closest.
It sounds silly and it changed something real. Engineers who had never opened the dashboard started opening it. Then they started asking which of the things in the sprint would actually move it, which is a question I had been asking them for months without much luck. Nobody made anyone care about the metric. They cared because they had money, or pride, on it.
That is the useful bit. The ritual is fun, and the fun is a delivery mechanism for alignment. You cannot instruct a team into caring about a number.
Most of the time, do not ship
I think seventy to eighty percent of features do not need to be shipped.
Most of a team's time should go on scalability, on buffer, on some genuine down time so creativity has room to happen, on the long threads that take a quarter to pay off, and on cleaning up the things we did badly the first time because we were in a hurry. That is the default state. Not a lot of visible output.
And then, when an insight genuinely works and it matters for growth, the team needs to move extremely fast. Not fast in the sense of a slightly tighter sprint. Fast in the sense that everything else stops.
The asymmetry is the point. You cannot sprint from a standing start if the team is already at ninety percent utilisation shipping features nobody needed, and you cannot sprint through a codebase you have been neglecting for a year. The slow months are what buy you the fast weeks.
This only works with a lot of trust in both directions. Engineering has to believe me when I say this one is P0, which means I have to be right most of the times I have said it before. If I call it wrong twice, nobody moves the third time, and the third one might be the one that mattered. Knowing what is actually P0, and having the credibility to declare it, is most of the job.
PMs are engineers and engineers are PMs
A PM is an engineer who happens to have slightly better customer insight. An engineer is a PM who happens to understand how to build a system that does not fall over.
I mean that fairly literally. The skills should be interoperable. A PM who cannot read the code is guessing about feasibility. An engineer who has never sat on a customer call is guessing about what matters. Both guesses are expensive.
What you actually get from pushing on this is empathy, which sounds soft and is not. When an engineer has watched a user fail at something, they stop arguing about whether it is worth fixing. When a PM has understood why a change is a three week job rather than a two day one, they stop asking for it on Friday.
Which twenty percent
The obvious objection to shipping twenty percent of features is that it sounds like a licence not to ship. Fair enough. The whole thing rests on being right about which twenty percent, and being right is not a talent. It is a workload.
So the actual job, most weeks, is validation. Before anything gets near a sprint I want to already know whether it works, and I want to have found that out without asking engineering for anything.
That last part is the constraint that matters. The moment validating an idea needs a build, you have spent the very thing you were trying to protect. So you do it some other way. Call twenty customers. Run the workflow by hand for a week and see if anybody notices when you stop. Fake the feature with a spreadsheet and a person doing the work behind it. Vibe code something ugly that nobody will ever see. Go and look properly at the data you already have instead of asking for new instrumentation.
Most ideas die in there. That is the point of doing it. By the time I am calling something P0, I am not making a bet, I am reporting a result. That is what makes it reasonable to ask a team to drop everything, and it is why they believe me when I ask.