Skip to content
NeuralYug

Work · Financial services · Rural Nepal

A voice-first Nepali helpline that answers members over an ordinary phone call

A blueprint for a member helpline that works in Nepali and Maithili over a normal phone call - fine-tuned speech recognition, answers grounded in the institution's own rules, and a warm handoff to staff the moment confidence drops.

blueprint10 Aug 2026
~55%
Share of Nepalis whose mother tongue is not Nepali (Census 2021)
11.05%
Maithili speakers as a share of population - the second language this blueprint targets (Census 2021)
~71%
Nepal literacy rate, 2021 - the remainder is why the channel is voice, not chat (World Bank / CBS)
124
Mother tongues recorded in Nepal, of which 21 cover ~95% of the population (Census 2021)

This is a solution blueprint — a reference architecture we can build for your business. The figures above are cited industry benchmarks for this class of system, not results claimed for a named client.

The pipeline we'd build

5 stages

Stage 01 · Call

A member calls the normal helpline number, or leaves a voice note - no app and no data plan required.

The challenge

A cooperative's call line carries the same handful of questions all day: what is my balance, when is the next instalment due, what documents does a loan need, why was a payment rejected. The questions are routine, but the callers are not a uniform group. Many speak Maithili or Bhojpuri at home rather than Nepali, a substantial minority cannot comfortably read, and a good number are calling from a feature phone with no data plan at all. A web portal or a text chatbot does not reach them - it reaches the subset that was already easiest to serve. Meanwhile staff spend their day reciting answers that already exist in writing, and the queue is longest exactly when field work is busiest.

What we did

The blueprint puts the assistant on the channel members already use: an ordinary phone call, plus a voice note on WhatsApp or Viber for anyone who prefers asynchronous. Speech recognition is fine-tuned on audio collected the way members actually call - regional accents, background noise, a mix of Nepali and Maithili in the same sentence - because published work on Nepali fine-tuning reports substantial word-error-rate reductions over off-the-shelf models, driven by exactly that kind of variation. Language identification runs alongside intent classification, so a caller is never locked to a single language setting. Answers are retrieved from the institution's own published rules and schedules rather than model memory, so a wrong answer cannot be invented and delivered confidently in a language the branch manager cannot audit. Every reply is spoken back through regional text-to-speech. When confidence drops, or the question touches money movement, the call is handed to a member of staff with the transcript already attached.

The outcome

The intended shape is narrow and measurable rather than sweeping. Three numbers define success and are agreed before any audio is collected: word error rate measured on the institution's own call recordings rather than a clean benchmark; containment, meaning the share of calls that reach a correct answer without a human; and the handoff rate, which should never be zero - a system that never escalates is either answering trivial questions or answering hard ones badly. The economic case rests on volume of repeat questions, not on replacing staff: the helpline absorbs the recitation, and staff keep the conversations that need judgement. Language coverage is deliberately staged - one language and one question type first, then the next - because breadth is what most often kills this class of project.

Stack

Fine-tuned Whisper-family ASRLanguage identificationIntent classificationRetrieval over institution rulesNepali / Maithili TTSSIP telephonyWhatsApp & Viber Business APIHuman handoff console

The reasoning behind this blueprint - why general models handle written Nepali well but stumble on speech and dialect - is set out in Your AI doesn't speak Nepali. This is what that argument looks like applied to one concrete, unglamorous use case.

Why a cooperative helpline, and not a chat widget?

Because the people a cooperative most needs to reach are precisely the ones a chat widget misses. Nepali is the mother tongue of 44.86% of the population; the remainder grew up with Maithili, Bhojpuri, Tharu, Tamang or one of another 119 languages. Literacy sat around 71% in 2021. A text interface quietly selects for the members who were already easiest to serve, and the helpline queue stays exactly as long as it was.

What does the call path look like?

Member call to spoken answer

The speech model and the handoff are where the difficulty concentrates. Everything else is ordinary infrastructure.

Reference architecture
The member's endUnderstandingAnswering

Tap any component above for its role and the real tech.

  1. Ordinary phone call (Client, Feature phone or basic Android): Reach is decided here. A call assumes neither a smartphone nor a data plan, which is why it is the primary channel rather than an app.
  2. Telephony gateway (Service, SIP trunk / WhatsApp & Viber Business API): Carries audio in and the spoken reply back. Ordinary infrastructure, but it sets per-minute cost and therefore the volume at which the helpline pays for itself.
  3. Speech recognition, fine-tuned (Model / AI, Whisper-family, tuned on member call audio): Tuned on recordings made the way members actually call. Published fine-tuning work on Nepali reports substantial word-error-rate reductions over off-the-shelf models, attributed to variation in speaker, acoustic environment and dialect.
  4. Language and intent routing (Service, Language ID + intent classifier): Decides both what is being asked and which language it is being asked in. Members code-switch mid-sentence, so language is treated as a per-utterance signal rather than a fixed per-member setting.
  5. Institution's own rules (Data, Retrieval over approved content): Balances, instalment dates, document lists and rate schedules come from approved sources. The assistant repeats what has been published rather than generating an answer, so it cannot invent a figure a branch manager would have to defend.
  6. Speech synthesis (Model / AI, Nepali / Maithili TTS): Speaks the answer back. Synthesis quality varies sharply by language, and it is usually what decides which languages can honestly be launched with.
  7. Handoff to staff (External, Warm transfer with transcript): Fires on low confidence and on anything touching money movement. The transcript travels with the call, so the member does not repeat themselves. How often this fires is the honest quality metric.

Tap any component for what it does and where this class of system usually fails.

Who does this actually reach?

Mother tongue, share of population

Two languages cover a majority of callers; the tail is long and real

Census 2021
011223445NepaliMaithiliBhojpuriTharuTamang

National Population and Housing Census 2021. Launching with Nepali and Maithili covers a majority of callers; each language after that is a separate decision with its own data and synthesis requirements.

Mother tongue, share of population — data table
CategoryShare of population (%)
Nepali44.9
Maithili11.1
Bhojpuri6.2
Tharu5.9
Tamang4.9

Language coverage is staged deliberately - one language and one question type at a time.

What would be measured?

Three numbers agreed before any audio is collected

Voice projects drift when success is defined after the build. These are written down first.

Acceptance criteria
MeasureWhy it decides the projectThe trap
Word error rate on the institution's own recordingsAccuracy on real calls - accent, noise, code-switching - is what members experience.Reporting it on clean read speech, where any model looks good.
ContainmentThe share of calls answered correctly without a human; this is where the economics live.Quietly redefining it mid-project to include calls that were abandoned.
Handoff rateTracks honesty. Falling steadily as the audio set grows is evidence the work compounds.Treating zero as the goal - a system that never escalates is answering hard questions badly.

Where this fits alongside our other blueprints

The grounded-answer pattern here is the same one behind the citizen-services assistant, and the staged rollout mirrors the crop advisory blueprint. The difference is the channel: those assume a screen, this assumes a phone call and a caller who may not read. If you want the reasoning rather than the architecture, start with the post, or see how we scope this kind of work under Neural AI.

At a glance

Client

Cooperative / microfinance

Sector

Financial services · Rural Nepal

Service

Neural AI

Kind

blueprint

Headline result

~55% · Share of Nepalis whose mother tongue is not Nepali (Census 2021)

Handover

Documented, tested code in your repository

Questions we were asked

Does this need members to have a smartphone?

No, and that is the point of the design. The primary channel is an ordinary voice call over SIP telephony, which works on a feature phone with no data plan. WhatsApp and Viber voice notes are offered as an additional asynchronous option for members who already use them, not as a requirement.

Which languages would it support?

The blueprint starts with Nepali and Maithili, the two largest mother tongues in Nepal at 44.86% and 11.05% of the population in the 2021 census. Further languages are added one at a time, gated on whether usable speech recognition and synthesis exist for them. Launching a half-working language is treated as worse than not offering it.

How does it avoid giving members wrong financial information?

Answers are retrieved from the institution's own published rules, schedules and rates rather than generated from model memory, so the assistant can only repeat what has been approved. Anything touching money movement, or any question where confidence is low, is routed to a member of staff with the transcript attached rather than answered.

Are these results NeuralYug has delivered?

No. This is a solution blueprint - a reference architecture that can be built. The figures are Nepal census statistics and published research benchmarks for this class of system, not outcomes claimed for a named client.

Related service

Neural AI

Agents and assistants that understand your business.

Same problem, different business?

We'll send the architecture and a realistic timeline for your version of this — no obligation.

Request a blueprint