All articles
AIEngineeringVoice AI

Our newest teammate is a bot. It understands Hinglish, remembers what we decided, and is always first on the call.

Tanveer Kaur
Tanveer Kaur

AI Engineer Intern

October 7, 20267 min read
Our newest teammate is a bot. It understands Hinglish, remembers what we decided, and is always first on the call.

How we built CosX Voice, and what it took to get a machine to understand “theek hai, so the heartbeat is failing, matlab we need logs.”

Speaker 1 · resolved: theek hai, so the charger heartbeat is failing, matlab we need the logs before Friday.

The short version

  • We built our own meeting notetaker (think Otter or Fireflies) for meetings that switch between Hindi and English mid-sentence.
  • Recording the call was easy. Getting the spelling right was not. Speech models write English words in Hindi script, and once that happens the original spelling is lost.
  • Then we added an AI you can ask “what did we decide?” about any past meeting. We also wrote tests that fail it whenever it gives a vague answer.

Key numbers

text
Number  What it means
------  -------------------------------------------------------------------------------------
~10x    faster than real time: a two-hour meeting transcribed in about ten minutes on one GPU
1 in 3  words in one two-hour meeting had to be rescued from Hindi script
1 day   to build the AI that answers questions across meetings
0       “who talked the most” leaderboards. On purpose.

This is a normal sentence in our office

Sit in on any CosX meeting for five minutes and you’ll hear it. Someone starts a sentence in Hindi, throws in a few English tech words, and finishes it in whichever language comes to mind first. “Theek hai, so the charger heartbeat is failing, matlab we need logs.” Nobody even notices. That’s just how we talk.

Software notices, though. Most speech tools expect one language per call, so they end up catching about half of every sentence, which doesn’t make for useful notes. We wanted something that could sit in our calls, follow how we actually speak, run on our own servers, and tell us a week later what we’d agreed to.

So we built our own and called it CosX Voice.


Meet the bot

Paste a Google Meet link into CosX Voice and a bot starts a real Chrome browser on a virtual screen, joins the call, and waits in the lobby like everyone else until someone lets it in. Then it records the audio and, four times a second, checks Meet’s own “who’s talking” highlight, so it can put a name next to every voice later.

Captions look like an easier way to tell who’s speaking, but they arrive a second or two after the words they describe. So the bot trusts the speaker highlight first and only falls back to captions when the highlight can’t decide.

Then we hit our favourite bug so far. The bot was set to leave if it had been alone in a call for 60 seconds, which seemed reasonable. The problem is that the bot joins on time and people don’t. So it kept showing up first, waiting a minute, deciding the meeting was over, and leaving before anyone else arrived.

Our bot was so punctual it kept leaving meetings before they started.

Now the 60-second rule only kicks in once at least one person has joined. We also added a three-hour limit, because no meeting should run longer than that (ours included).


One word broke our transcripts

Speech models like Whisper are actually pretty good at Hindi, and oddly enough, that was the problem. If an English word shows up in a Hindi sentence, the model writes it out by sound in Devanagari instead of leaving it in English. “Notification” comes out as नोटिफिकेशन.

Doesn’t seem like a big deal, until you convert it back into English letters so everyone on the team can read it. Go letter by letter and you get notiphikeshan. No rule can turn “keshan” back into “-cation”, because the original spelling was already lost in the step before.

text
What was said        ->   What the model wrote   ->   Turned back into English
notification         ->   नोटिफिकेशन              ->   notiphikeshan

Hinglish speech, run through a speech model
CosX Voice, field notes

And it happens a lot. In one two-hour meeting, 5,656 of the 18,537 words had to be converted out of Hindi script, close to a third of everything said. In calls with more Hindi, it was over half.

The Hinglish decoder

Real examples from our transcription experiments. Each word goes through the same four stages: what was actually said, what the model wrote, what a letter-by-letter conversion gives you, and what it should have said.

notification

  1. What was said: notification. An English word, said in the middle of a Hindi sentence.
  2. What the model wrote: नोटिफिकेशन. Written by sound, in Devanagari.
  3. Converted letter by letter: notiphikeshan. The “-cation” is gone. No rule can bring it back.
  4. What it should say: notification. Only fixable by having the model write English in the first place.

karna

  1. What was said: karna. “To do”. Possibly the most common verb in any meeting.
  2. What the model wrote: करना. Correct Hindi. So far so good.
  3. A Sanskrit-style library: karanA. Keeps every vowel, the way Sanskrit does. Hindi doesn’t.
  4. With our rules: karna. Drop the vowel between two syllables.

dekh

  1. What was said: dekh. “Look”, as in “dekh lete hain”, let’s take a look.
  2. What the model wrote: देख. Correct Hindi.
  3. A Sanskrit-style library: dekha. Adds a vowel nobody says.
  4. With our rules: dekh. The vowel at the end of a word always drops.

namaste

  1. What was said: namaste. The first word of half our client calls.
  2. What the model wrote: नमस्ते. Correct Hindi.
  3. Our rule, applied blindly: namste. Drop the middle vowel everywhere and you get a word nobody can say.
  4. With the exception: namaste. Never drop it before a cluster of consonants.

jobs

  1. What was said: jobs. An English word, in a Hindi sentence.
  2. What the model wrote: जॉब्स. Uses ॉ, the vowel Hindi borrowed for English words.
  3. Every library we tried: jaॉbsa. None of them knew ॉ, so the word came out half converted.
  4. What it should say: jobs. Same story for doctor, college, and most of a tech meeting.

We asked for English and got Korean

Whisper lets you pass in a prompt, a bit of sample text it uses as a style guide. The obvious idea was to give it a prompt with Hindi and English mixed together, so it would learn to leave English words alone. Everyone we asked said the same thing: just prompt it.

So we tested it. We took 90 seconds of real meeting audio, ran it with four different prompts, and counted how many English letters came back.

More English in, less English out

text
Prompt we gave the model                             English letters in the output
---------------------------------------------------  -----------------------------
Hindi script only                                    0
A light mix (with one English phrase planted in it)  13
Heavy Hinglish (lots of English mixed in)            7, worse
Hindi written fully in English letters               95, the model fell apart

90 seconds of real meeting audio, Whisper large-v3-turbo.

The 13 letters from the light mix spelled “cardiac arrest”, which was just the phrase we’d put in the prompt, copied back to us. Adding more English actually made the output less English. The last prompt was the worst of all: the model came back with bits of Korean and Spanish, from a meeting that was entirely in Hindi and English.

What we took away from it: the prompt only nudges the vocabulary. The output language comes from one setting, the language you pass in. Once we accepted that, things got a lot easier. Setting the language to English worked whether a clip started in Hindi or in English. The “translate” mode, which sounds like the right choice, dropped a whole “how are you doing” from one clip and repeated another phrase three times.

The libraries were made for Sanskrit

Next we tried off-the-shelf transliteration libraries. Most of them are built for Sanskrit, where every syllable keeps its vowel. Spoken Hindi drops a lot of those vowels. So देख came out as “dekha” instead of “dekh”, करना as “karanA” instead of “karna”, and जॉब्स came out as jaॉbsa, half converted. None of them handled ॉ, the vowel Hindi picked up for English words like job, doctor and college. A tech meeting is full of words like that.

In the end we wrote our own rules. Drop the vowel at the end of a word (baat, never baata). Drop it between two syllables (karna, not karana). But never before a cluster of consonants, or namaste turns into “namste”. We also wrote all of this down with the numbers, so the next time someone says “let’s just prompt it”, we have something to show them.


Nobody rereads a transcript

Let’s be honest, though. That two-hour meeting produced 18,537 words, and nobody is going back to read them. What people actually want to know a week later is “what did we decide about tariffs?” or “who said they’d pull the logs?”

So we built an AI agent that reads the transcripts and answers those questions. Ask it “What came out of our latest meeting?” or “Where are we on charge-point diagnostics?” and it catches you up the way a colleague would. You get a short summary of where things stand, pulled from every meeting where the topic came up, followed by what was agreed, who’s doing what, and what’s still open. If something was never discussed, it tells you that rather than guessing.

text
How a meeting becomes an answer

  Google Meet          the call, in Hinglish
       |
  CosX Voice bot       joins, records, tracks who's speaking
       |
  Transcript           speaker-labelled and readable
       |
  AI notes             decisions, owners, open questions
       |
  "What did we decide?"   answered across every meeting

Ask it “Where are we on charge-point diagnostics?” and the answer comes back in four parts:

  • The story so far, drawn from three meetings
  • Agreed
  • Who is doing what
  • Still open

Building the agent took one day. More on why further down.


How we know it isn’t making things up

1. It only sees what it needs

We don’t give the AI timestamps, call lengths or how long each person spoke. Nobody wants their weekly standup turning into a “who talked the most” leaderboard, and if the AI never has those numbers, it can’t share them. It also won’t paste the raw transcript into its answers unless you ask for a quote.

2. Tests that vague answers can’t pass

Each test question checks for specific facts, sometimes from meetings the question doesn’t even mention. On top of that, a separate script feeds every test a vague but confident-sounding answer and makes sure the test fails it. If a generic answer could pass, the test wouldn’t be worth much.


Things our AI had to learn the hard way

One five-minute test recording was transcribed ten times

Left alone, that would make it the most talked-about meeting in our history. The AI now knows to skip it.

Two meetings, one name, one date

The AI doesn’t ask “which one?”. It answers for both and tells them apart by what each covered.

Two teammates whose first names sound identical

Different people with different action items. The AI didn’t work that out on its own. We had to tell it.


Why it only took a day

Honestly, this wasn’t our first go. It’s the fifth AI we’ve built that answers questions over a company’s own data. When we finally put the first four next to each other, it was a bit embarrassing.

text
What we found    Where
---------------  -------------------------------------------------------------------------------
byte-identical   one core file, across two projects for different clients on different databases
667 · 667 · 664  lines in the API file, project after project
8 tools          the same core set every time, with the same dependencies

We’d basically been copy-pasting the same agent from project to project, and whenever we fixed a bug in one, the others never got the fix. So we pulled the shared parts out into one framework, with 542 tests behind it. Now a new agent is mostly a description of the data, a few rules in plain English, and a couple of custom tools. The meeting agent needed just 242 lines of its own code, and that’s how it got done in a day.


Voice AI for the way India actually talks

At CosX we build voice bots and AI agents that run on your own data and work in the languages your team and customers actually speak. If you’re building something with voice, or you have a pile of meeting recordings, call logs or documents nobody has time to go through, we’d be happy to talk.

Also, we’re curious: what’s the weirdest thing your notetaker has ever written down? Ours gave us “notiphikeshan”.

Work with us

Have a project in mind? Let's talk.

Pilots, platforms, or roadmaps — tell us what you're building and we'll get back within one business day.

Newsletter

Get our latest writing in your inbox.

Agentic engineering, AI platforms, and what we learn shipping them — no spam, unsubscribe anytime.