Trading agents · Round 1 result
Arnav Chauhan won with +5.91%.
QQQ ended the same window at -3.9%. Arnav's code is public and now runs as the benchmark every Round 2 entrant must beat.
Bring one measurable problem. We build the test, builders compete under published rules and scoring, and you pay only when a solution clears the agreed bar.
Public challenge boards
How it works
01
Tell us the input, the useful output, and what solving it is worth.
02
We agree the data, scoring, qualifying bar, and practical limits before launch.
03
Builders compete on the same evaluation. Only work that clears the bar can win.
If nothing clears the bar, there is no forced winner. The bounty rolls over or comes back.
For problem owners
Bring a workflow, footage, or dataset that matters. Builders return an answer, evidence, and run cost against the same test. Kitchen CCTV is one live example.
Challenges
Enter a live contest, inspect a finished result, or help shape the data for the next one.
For restaurants and cloud kitchens: answer hidden questions from CCTV-style footage about caps, handoff, bottlenecks, visible steps, and what was not visible.
4 scored entries
Write one Python function that beats Arnav on a fresh live market window — no finance background, no money, no API key.
Live participant board
Turn official filings and public sources into verified profiles for every Norway-registered company—and keep them current.
$2,500 prize · partner with Håvard to launch in Norway
Trading agents · Round 1 result
QQQ ended the same window at -3.9%. Arnav's code is public and now runs as the benchmark every Round 2 entrant must beat.
Local dictation · shipped as RambleFix
RambleFix types English and mixed Hindi+English at the cursor, runs locally on your Mac, and costs nothing. It was built specifically for code-switched speech.
We are shaping the test data and scoring now. Share representative data or register interest; the bounty and rules will be published before entries open.
Turn sports footage or event feeds into accurate scores and game-state updates.
What we need: Sample footage or an event feed, plus trusted final scores and game states.
Estimate the odds of real-world events, explain the signal, and trade against a fixed benchmark.
What we need: Real questions, source odds, decision dates, and unambiguous outcomes.
Match purchase orders, invoices, and delivery records, then surface the exceptions that need action.
What we need: Anonymized purchase orders, invoices, delivery records, and resolved exceptions.
Have a problem? Post it as a challenge and we will turn it into a scored test.
Live boards
For restaurants and cloud kitchens: answer hidden questions from CCTV-style footage about caps, handoff, bottlenecks, visible steps, and what was not visible.
6 submissions reviewed; 4 have verified scores. Answer accuracy ranks first, then lower model/API cost.
| # | Entry | Answers correct | Model cost / 60 min |
|---|---|---|---|
| 1 | Kim Beomjin · verified runAug 26 rerun on commit df0f143 completed cleanly in 43 seconds. It used 11 frames and 6 local model calls; the person count at 00:45 and island interaction at 15:00 remain the misses. | 4 / 6 | $0.00 |
| 2 | Vishal · verified runThe Aug 26 rerun on commit 546bca3 improved to 4 / 6. It completed cleanly in 4m 59s at $0.00; head-cover counting and the 00:45 island-contact check remain the misses. | 4 / 6 | $0.00 |
| 3 | Meet · verified runTied on answers; the Aug 10 cloud run stayed under the cap at $0.0312 per 60 minutes. | 3 / 6 | $0.0312 |
| 4 | Raam · verified runAug 20 rerun on commit 2854a8c completed cleanly against the official six-question set at $0.00. The pipeline returned valid answers and evidence; person counts and action detection remain the main gaps. | 2 / 6 | $0.00 |
Aug 26 rerun on commit df0f143 completed cleanly in 43 seconds. It used 11 frames and 6 local model calls; the person count at 00:45 and island interaction at 15:00 remain the misses.
The Aug 26 rerun on commit 546bca3 improved to 4 / 6. It completed cleanly in 4m 59s at $0.00; head-cover counting and the 00:45 island-contact check remain the misses.
Tied on answers; the Aug 10 cloud run stayed under the cap at $0.0312 per 60 minutes.
Aug 20 rerun on commit 2854a8c completed cleanly against the official six-question set at $0.00. The pipeline returned valid answers and evidence; person counts and action detection remain the main gaps.
No score is published until setup and integrity checks pass.
The Aug 26 rerun on commit b501a60 completed within the frame and runtime caps, but benchmark-specific timestamp answer overrides failed the generalization check. No verified score is published.
Commit eebd656 compiled, but the evaluator could not complete the declared 4B model setup, so no official answer file or score was produced. A reproducible model warm-up path is required.
Watch · 2 minutes
A problem you keep meaning to deal with — and the move most people don't have a habit for yet: turn it into a challenge, let builders compete on the same test, and give qualifying builders a public proof of what they built.
After a published result, builders can claim their profile. Qualified builders who want a signed certificate can share that profile on LinkedIn with what they built and liked or learned; sharing never changes the score. Browse public results →
Public proof
These are public examples of builders meeting a measurable bar. We publish the shared inputs, the score, the limits, and the work so the result can be checked.
Trading · published graph

Diagnostic graph from the Jul 7–27 live window; the round remained open.
Open the winner project →Local dictation · shipped outcome
Sankeerth's winning approach shaped how RambleFix remembers important words. Arnav's work helped improve product and technical terms in Hindi+English mode.
RambleFix types English and mixed Hindi+English at the cursor, keeps audio on your Mac, and costs nothing. It was built specifically for code-switched speech—not just a generic multilingual language list.
See the builders, code, and measured results behind the public challenges.
Browse public results →Bring one AI or automation problem with an answer we can measure. We define the test, recruit builders, and run the competition. Here are a few examples:
If you can describe what goes in and what good looks like, we can usually turn it into a test in about a week, then run a few weeks of competition. The bounty is paid only when a solution clears the bar you approved.
How much should the bounty be? Size it to the problem. A simple way to think about it: if cracking it is worth a builder's weekend, the prize should feel worth that weekend. Around $1,000+ can work for a focused task; larger pools pull in stronger fields.
Send your problem →Two friends — product-and-tech geeks, both ex-founders — who hit this exact problem in our own work every week: you need an agent for something that matters, and no honest way to tell which approach is actually best. builderr is our fix.
We put real money on it. Each challenge has its own sponsor and a published test. We show what was run, what cleared the bar, and what builders can inspect afterward. No black box: the point is proof, not promises.
Current challenges have no platform fee. The bounty passes through to qualifying winners, minus agreed running costs. If no solution clears the bar, there is no forced winner. Have a problem we can score? Get in touch.