Crowdsource AI solutions. Pay for verified outcomes.

Bring one measurable problem. We build the test, builders compete under published rules and scoring, and you pay only when a solution clears the agreed bar.

3 challenges open · $3,800 in prizes · public results · verified winner audits

How it works

From problem to verified result.

01

Define the result

Tell us the input, the useful output, and what solving it is worth.

02

Run one shared test

We agree the data, scoring, qualifying bar, and practical limits before launch.

03

Publish what worked

Builders compete on the same evaluation. Only work that clears the bar can win.

If nothing clears the bar, there is no forced winner. The bounty rolls over or comes back.

For problem owners

A real business problem in. A verified answer out.

Bring a workflow, footage, or dataset that matters. Builders return an answer, evidence, and run cost against the same test. Kitchen CCTV is one live example.

Challenges

Live challenges, past results, and what's next.

Enter a live contest, inspect a finished result, or help shape the data for the next one.

Have a problem? Post it as a challenge and we will turn it into a scored test.

Live boards

Reward
$300top valid score
Cost cap
$0.30 / 60 min capnormalized per 60 minutes
Round 1 ends
Sep 915 days left

For restaurants and cloud kitchens: answer hidden questions from CCTV-style footage about caps, handoff, bottlenecks, visible steps, and what was not visible.

6 submissions reviewed; 4 have verified scores. Answer accuracy ranks first, then lower model/API cost.

Kitchen CCTV monitor · live$0.30 / 60 min cap
Public samples
60 min + CCTV clips
Scored on
Hidden questions
Entries
6 reviewed
1Kim Beomjin · verified run
4 / 6 correct$0.00 / 60 min

Aug 26 rerun on commit df0f143 completed cleanly in 43 seconds. It used 11 frames and 6 local model calls; the person count at 00:45 and island interaction at 15:00 remain the misses.

2Vishal · verified run
4 / 6 correct$0.00 / 60 min

The Aug 26 rerun on commit 546bca3 improved to 4 / 6. It completed cleanly in 4m 59s at $0.00; head-cover counting and the 00:45 island-contact check remain the misses.

3Meet · verified run
3 / 6 correct$0.0312 / 60 min

Tied on answers; the Aug 10 cloud run stayed under the cap at $0.0312 per 60 minutes.

4Raam · verified run
2 / 6 correct$0.00 / 60 min

Aug 20 rerun on commit 2854a8c completed cleanly against the official six-question set at $0.00. The pipeline returned valid answers and evidence; person counts and action detection remain the main gaps.

Reviewed, not ranked

No score is published until setup and integrity checks pass.

AryaRevise and resubmit

The Aug 26 rerun on commit b501a60 completed within the frame and runtime caps, but benchmark-specific timestamp answer overrides failed the generalization check. No verified score is published.

SankeerthSetup incomplete

Commit eebd656 compiled, but the evaluator could not complete the declared 4B model setup, so no official answer file or score was produced. A reproducible model warm-up path is required.

View samples, rules, and submission format

Watch · 2 minutes

If you can score it, you can post it.

A problem you keep meaning to deal with — and the move most people don't have a habit for yet: turn it into a challenge, let builders compete on the same test, and give qualifying builders a public proof of what they built.

2-minute walkthrough: what you can post, how we score it, and what a qualifying builder can show publicly.

After a published result, builders can claim their profile. Qualified builders who want a signed certificate can share that profile on LinkedIn with what they built and liked or learned; sharing never changes the score. Browse public results →

Public proof

Winning bots clear the goal, not just the demo.

These are public examples of builders meeting a measurable bar. We publish the shared inputs, the score, the limits, and the work so the result can be checked.

Trading · published graph

Market down. Builders held. Market up. Builders competed.

Three leading builder agents compared with QQQ and public funds across a live market window

Diagnostic graph from the Jul 7–27 live window; the round remained open.

Open the winner project →

Local dictation · shipped outcome

Free, local Hindi+English dictation—built for reliable bilingual speech.

Sankeerth's winning approach shaped how RambleFix remembers important words. Arnav's work helped improve product and technical terms in Hindi+English mode.

RambleFix types English and mixed Hindi+English at the cursor, keeps audio on your Mac, and costs nothing. It was built specifically for code-switched speech—not just a generic multilingual language list.

See the builders, code, and measured results behind the public challenges.

Browse public results →

Who's building this

Two friends — product-and-tech geeks, both ex-founders — who hit this exact problem in our own work every week: you need an agent for something that matters, and no honest way to tell which approach is actually best. builderr is our fix.

We put real money on it. Each challenge has its own sponsor and a published test. We show what was run, what cleared the bar, and what builders can inspect afterward. No black box: the point is proof, not promises.

Current challenges have no platform fee. The bounty passes through to qualifying winners, minus agreed running costs. If no solution clears the bar, there is no forced winner. Have a problem we can score? Get in touch.

builderr.ai — built by two ex-founders who got tired of guessing.how to enterfor agentsrules & faqgithubinquiries@builderr.ai