Picture this. Somewhere in India, a person fills out a form at 11 pm because they have finally decided to do the thing they had been putting off for months. Learn the skill. Buy the policy. Book the demo. For about ninety seconds, they are the easiest customer that business will ever have.

Nobody calls them back until Thursday.

By Thursday they are a different person. The urgency has cooled. Someone else has called first. The business logs it as a bad lead, and it never finds out that it was a good lead with a bad response time.

This is not some rare edge case. It happens across Indian businesses every single day, and almost nobody treats it as a technology problem. It gets treated as a staffing problem, which is why the answer is always hire more callers — and why the answer always runs out.

That gap is the entire reason DialNexa exists.

The arithmetic nobody escapes

Every business we talk to already knows what it should be doing. Call every lead in the first five minutes. Follow up three times, not once. Confirm every appointment the day before. Chase the payment before it ages into a bad debt. Call back the customer who raised a ticket instead of making them wait on hold.

Nobody disputes any of this. They just can’t do it, because a team of twelve people can make a fixed number of calls in a day, and that number is far smaller than the number of calls the business actually needs. So the work gets triaged. The top of the list gets called. The rest of the list quietly becomes the reason the quarter missed.

The interesting thing about that constraint is that it has nothing to do with intelligence or effort. It is arithmetic. And you do not fix arithmetic with another hiring plan. It is exactly the kind of problem a machine should take off a person’s hands.

We are not building this so that businesses can talk to their customers less. We are building it so that the calls that were never going to get made, get made.

Why voice, and why now

For years, voice AI made for a great demo and a terrible colleague. It could survive the happy path — a clear question, a clean line, a patient caller — and then fall apart the moment a real person interrupted, changed languages, or answered a question with another question. Businesses were right not to trust it with the calls that mattered.

That boundary has moved.

Speech recognition now holds up in conversations that would have broken it a few years ago. Language models can follow intent instead of forcing every caller through a decision tree. Modern telephony and inference infrastructure can put all of that on a live call quickly enough that the person on the other end does not have to slow down for the machine.

The important change is not that AI can speak. It is that, for the first time, it can listen, reason and act inside the narrow rhythm of a phone call — while the reason for the call still matters.

At the same time, the cost of not calling has become impossible to ignore. Customers expect an answer now. The best lead is often the one that arrived after the team went home. The businesses growing fastest cannot keep matching every spike in demand with another hiring cycle. Voice is no longer a futuristic interface looking for a job. It is becoming infrastructure for work that already exists and is already being left undone.

This is the moment we have been building for: the shift from voice AI that sounds impressive for two minutes to voice AI that can be trusted for two million calls. When we started DialNexa, that was the bar we set for ourselves.

We started with the hard version on purpose

If you wanted an easy demo, you would build for one language, in a quiet room, on good wifi, with a caller who speaks in complete sentences.

India gives you none of that comfort.

A single conversation here will start in English, pivot to Hindi at the point where the caller gets emotional, and land on a number spoken in a way that no dictionary contains. The next call is in Marathi. The one after that is a Tamil speaker in Coimbatore whose name most speech systems will mangle on the first try. Names, places, amounts — the exact tokens that matter most in a business call — are the ones a system trained somewhere else gets wrong.

This is why we support 70+ languages, and why Indian regional languages and Hinglish are not a feature we added later. Hinglish in particular is not a language you can bolt on, because it isn’t a language — it is two languages sharing a sentence, switching at the clause, sometimes mid-word. You either build for that from the beginning or you build something that fails at exactly the moment the caller stops being polite and starts being real.

We chose the hard version first because a system that survives an Indian sales call at 6pm on a patchy network will survive anything. The reverse is not true.

Three hundred milliseconds

If we could put one number on the front of the product, it would be this one: sub-300ms.

That is how long the agent takes to start responding. It sounds like a specification, and it is really a design philosophy, because that threshold is roughly where a human stops experiencing a pause and starts experiencing a conversation.

Above it, everything goes wrong in a specific and familiar way. The caller thinks the line dropped. They say “hello?” The system, still processing, talks over them. Both back off. Both start again. Within two turns the person on the other end knows they are talking to a machine, and the moment they know that, they stop telling you anything useful.

Below it, something different happens. People interrupt. They change their mind halfway through a sentence. They ask the question they actually had rather than the question the menu allowed. They behave like themselves.

Everything we care about technically — how fast we transcribe, how early we commit to a response, how we handle a barge-in — is downstream of protecting that number. It is not a benchmark we publish. It is the thing that decides whether a call is a conversation or an obstacle.

What we actually believe

A voice agent should be judged on outcomes, not on transcripts. Did the meeting get booked. Did the payment get made. Did the customer get their answer without being transferred twice. Everything else is a demo.

The unglamorous calls matter most. Appointment confirmations, renewal reminders, document follow-ups, collections. Nobody puts these on a conference slide. They are where the money quietly leaks out of a business, and they are the calls that never get made.

Being multilingual is table stakes, not a differentiator. In India, a system that only works in English does not work.

A phone call is a promise. Someone gave you their number expecting a person’s attention. If we take that call and make it worse, we have not automated anything — we have just found a cheaper way to disappoint people. That constraint should be uncomfortable, and we intend to keep it that way.

Where we are headed

The near term is unglamorous and specific: make the conversations better. Fewer clarifying questions. Better handling of the moment a caller interrupts. More reliability on the tokens that actually matter — names, amounts, dates, addresses. Deeper hooks into the systems where the work actually lands, so that a booked meeting is a calendar entry and a captured objection is a CRM field, not a transcript someone has to read later.

The further-out version is a bigger claim, and we will say it plainly: we think every business conversation that should happen, will happen. Not most of them. Not the top of the list. The whole list — in the caller’s language, at the moment it is useful rather than three days later, at a cost that makes calling a customer back an obvious decision instead of a budget question.

We are a long way from that. But it is the right thing to be a long way from.

Why this newsletter exists

Voice AI is moving faster than anyone can read about, and most of the coverage is either a launch announcement or a thread claiming everything changed this week. Very little of it is useful to someone who has to actually put an agent on a real call on Monday.

So every week, this will be short and specific: what genuinely moved in Voice AI, what we shipped, and the research worth your time — including the findings that are inconvenient for us. When we get something wrong, we would rather say so here than have you discover it on a call.

That is the whole promise. I will keep it useful, I will keep it honest, and I will write it in the same plain language in which we discuss these problems inside DialNexa.


If you want to see what the arithmetic looks like on your own numbers, start with the product. And if you disagree with something in here — especially the 300ms claim — reply and tell us. That is a better use of this list than another launch announcement.

— Aditya