Results publishedlocal dictation

Sankeerth won the $500 dictation challenge.

The final six-clip review tested English and mixed Hindi-English speech on a local computer. Sankeerth scored 69.92 out of 100. This round is closed.
Prize
$500 to the winner
Completed
Aug 5
Scored on
Text accuracy + speed
Test
Six Hindi+English clips
  1. 1BuildUse the starter brief.
  2. 2SubmitSend one version of your code.
  3. 3TestWe run the same published test.
  4. 4ScoreThe metric below sets your rank.

Entries are closed. Use the brief to study the task and results.Download starter brief ↓

About these results: The result covers one six-clip Hindi+English review, not general dictation performance.

The round is closed. You can still study the plain-language starter brief, read the technical getting-started guide, fork the template, and inspect the final code and scoring notes.

How the result was scored

The $500 prize was decided by one combined score: 70 points for the final transcript and 30 points for how quickly that final is ready to paste.

Words and language. Does it write what was actually said, without translating the Hindi away? The final review checked meaning, names, numbers and other important facts.

Speed of the final. How fast a clean, faithful final lands after you stop talking — the target is ~2 seconds. You don't build a server.Our harness plays real speech into your engine at real time (one simple function — no networking) and measures when a stable final comes back. What's scored is that final — not the live preview while you're still talking; whether you reach it by streaming or batch is your call.

Only the final counts. Early drafts, first-word speed, and rewrites while you are speaking receive no points. A blank or hung final fails that clip.

Everything is scored offline on a single pinned machine — an Apple M1 Pro (8-core CPU + Neural Engine), 32 GB, macOS, accelerator on, network blocked — the same box for every entry and for the RambleFix benchmark. Same hardware, warmed up, identical conditions for comparing final-text speed. The exact latency targets and weights are in the public streaming contract.

Study the challenge and run the examples
2 minutes: why builders entered — and what they built.

The task

If you build with AI, you talk to it all day — and typing is the slow part. The useful target is a Wispr Flow-style dictation tool that does not send your speech to the cloud. It should work on a normal laptop, be private by default, and handle mixed-language speech.

The round focused on English and mixed Hindi-English clips. The published result covers that six-clip review; it does not establish performance on all speakers or recordings.

What the round required

  1. Get plain English right — the leading submissions already beat the open-source baselines we tested on sample clips.
  2. Land the final fast — the finished text lands within about 2 seconds of you stopping. That's the product feel to match, except local and offline.
  3. Keep both Hindi and English — write what was actually said (don't translate it to English). This is the real test.
  4. Stay on the laptop — no internet while it's scored.
  5. Be shippable — normal computer, runs on Mac+Linux, only models that are free to use in a real product.
  6. Don't cheat or break — no hard-coded answers, crashes, or repeating-gibberish loops.

Both languages mattered — the Hindi+English clips are the gate. The task required preserving both languages rather than translating everything into English. The exact rules and reference engine are in the reference-bot write-up.

Study the benchmark and build on it

RambleFix is open. Pull it, run it on your own machine, see exactly where it's weak, and build on top. It uses two models racing — a fast draft and a faithful final — both free-to-ship, no API keys, fully local.

Local setup example
# draft path — whisper.cpp small, loopback only
whisper-server -m ggml-small.bin --host 127.0.0.1 --port 8089 &

# faithful final — a public Hinglish model (Apple GPU)
export RAMBLEFIX_WHISPER_CPP_SERVER="http://127.0.0.1:8089/inference"
export RAMBLEFIX_HINGLISH_FINALIZER="python baseline/finalizer_hinglish.py"

Full steps + a ~50-line public finalizer (Oriserve / qwen3-hinglish) are in the baseline folder. The same draft() contract you implement.

Where the benchmark was weak:

  • RambleFix is strong, but not unbeatable. The most direct path is a faster final that still keeps both Hindi and English.
  • Hindi-heavy and fast switching clips were the risk. That is where the final review separated the field.
  • It drops some terms and numbers. Preserving names and numbers mattered to the final-text score; see the linked scoring contract for the exact rules.

Get the code →

How entries were reviewed

  • Read the getting-started guide and the build skill (the high-level architecture to follow), fork the template, and test on the included sample clips (English + Hindi+English) with python preview.py — all offline.
  • For the speed score, your engine just needs to return a fast final after the audio stops — one simple function, the shape's in the streaming contract; no servers, we drive the timing and feed the audio in real time.
  • Declare your models and their licenses — they must be commercial-friendly, so the winning tool can actually be released for free.
  • For the closed round, the pinned scoring machine ran the final verified review offline (Apple M1 Pro · macOS · network blocked). The published board keeps each entrant's last clean result.

Who's behind this challenge

This challenge is sponsored by Amit — he backed the $500 bounty and built the RambleFix benchmark. The challenge work now feeds a real free, local product: try RambleFix or open the source.

View final results →

Have a different problem to put a bounty on? Post a challenge →