Open for entriescompany research

Build an agent that finds company information online.

Give it a Norwegian company number. It finds public facts, links each fact to a source, and keeps the profile current.
Prize
$2,500 final rewards
Closes
Oct 21
Scored on
50 recall · 30 precision · 12 synthesis · 8 UX
Benchmark
Same checked company collection
  1. 1BuildUse the starter brief.
  2. 2SubmitSend one version of your code.
  3. 3TestWe run the same published test.
  4. 4ScoreThe metric below sets your rank.

Download the brief, run a sample, then follow the submission steps below.Download starter brief ↓

First run: one saved company, about 5 minutes. No API key or company download. This is practice, not an entry. Download the starter → View the sample site → Send a blocker →

01 · What to build

Find the information. Get it right.

Build a company research agent. Collect company details, people, locations, financial results, jobs and public activity where available.

Check each fact belongs to the right company. Link it to its source and date. Clearly mark anything you could not find.

Keep profiles current. Check sources again, show what changed and preserve the earlier evidence.

Your starting list: 411,160 eligible Norwegian companies with 2025 annual-account records. Your agent must work from a company number, including companies it has not researched before.

02 · What to submit

Submit your agent and a 100-company smoke test.

  • A 100-company smoke-test result or run report showing that your agent returns one complete result per input.
  • Your repository link, exact commit hash and one command to run the agent.
  • The models, APIs and licences you use, expected cost per official run, and contact details.

Use the public company list to test at any scale, including 1,000 companies or more. You do not need to submit precomputed profiles or a company manifest. Builderr supplies the companies for every official run.

Submit to Builderr →

03 · How we test it

The same official company batch for every eligible entry.

Builderr chooses the official company batch after the cutoff and gives the same batch to every eligible entry. The current shared set can grow over time—for example, from 1,000 to 1,100 companies. When it grows, every active entry is run on the same current set.

Your local 100-company smoke test is only for checking that your code works. It is not your competition score, and your precomputed profiles are not used for ranking. We score the agent on the fresh companies Builderr supplies.

Each official run has a fixed time and resource budget. Finish the run and return one result for every supplied company. If your agent times out or drops results, that run is not scored. If Builderr’s infrastructure fails, we rerun it.

04 · Common traps

Most failed runs fail here, not on the research.

Your agent picks its own companies. We hand it the official company batch at run time, chosen after the daily cutoff. Some will be companies you have never seen. Read them from that file. An agent working from its own list cannot be scored, however good the research is.

It handles one company at a time. We need one result per company in the batch, including when the official batch grows beyond 100 companies.

There is no single command to run it. Give us one command we can paste. A list of steps or a notebook cannot be run automatically.

It does not install on a clean machine. Clone your own repository into a new folder at the pinned commit, make an empty virtual environment, run only your declared install step, then your run command. That check matches ours and catches packages you installed months ago and forgot about.

It returns fewer than 100 results. Every company comes back, including the ones you found nothing for. See the result envelope below.

You think you cannot use an LLM. You can. Each run has a small external API budget, and if you need a model key for scoring, ask and we will supply one. What we cannot use is a credential tied to your own account on another service, because we cannot reproduce your run with it.

05 · How you score

Find more facts. Keep them accurate and current.

What to optimize: first make sure every fact belongs to the right company and has evidence. Then increase how much checked information you find. A high score cannot make up for a material wrong-company match.

50points

Recall and coverage: how much did you find?

We compare every agent with the same checked collection. 70% measures how many companies you covered and 30% measures how many checked facts you found.

30points

Is the information correct?

We check that each fact belongs to the right company and has a valid source and date. Wrong or unsupported facts lose points. A material wrong-company match also blocks qualification.

12points

Synthesis: is the result useful?

The profile should explain what the company does, what changed and what remains unknown, with sources for its conclusions.

8points

UX: is it easy to use and verify?

We check whether a user can find, compare and verify company information on desktop and mobile.

How do we know how much you found?

We combine findings from all submissions and Builderr’s own crawlers. We check the sources, remove duplicate facts and wrong-company matches, then compare everyone against that same collection.

For each type of information, 70% of its coverage score comes from how many companies you covered; 30% comes from how many individual facts you found.

Example: our checked collection has 50 job postings across 20 companies. You find 30 postings across 15 of those companies. That is 60% of postings and 75% of companies.

70% × 75% + 30% × 60% = 70.5% for this information type.

This feeds the 50 recall and coverage points, not your total score. Precision and evidence, synthesis, and UX make up the remaining 50 points. New verified findings update the collection and everyone’s scores. We finalize it after checking all eligible final submissions; it is not a claim that we found everything online.

06 · Qualification

Meet the score bar on an official run.

  • At least 65/100 overall on an official run.
  • Recall and coverage (50), precision and evidence (30), synthesis (12), and UX (8) make up the score. They are not separate qualification thresholds.
  • Return a result for every company Builderr supplies, including when information is missing or blocked. Never replace missing values with zero.
  • Keep the source, retrieval date and relevant reporting period for every claim. Preserve earlier evidence and avoid duplicate records when rerunning.
  • Unsafe source or secret handling, fabricated claims, compromised evidence, and evaluator failures are investigated with named ownership. A Builderr harness or shared-source failure is rerun and is not assigned to the builder.

Your score and qualification are separate.

A small factual mistake lowers your accuracy. A material wrong-company match is more serious because it can put an entire website, brand and set of facts under the wrong profile. Your score remains visible, but the run cannot become official until that match is corrected.

It is better to miss some information than publish it under the wrong company. When the company match is uncertain, return ambiguous or not available instead.

Return a result for every company.

We hand your agent a company batch and need one result per input, including the ones you found nothing for. Each result carries one of these states:

  • available — you found it
  • not_available — you looked, there is nothing there
  • blocked — the source refused the request
  • not_applicable — the question does not apply to this company
  • ambiguous — you could not be sure it is the right company
  • failed — the run broke on this one

If we hand you a batch and some results are missing, we cannot tell whether the rest were blocked, empty or crashed, so the run cannot be scored. “I found nothing” is a valid answer. A missing row is not.

Every official run: finish within the supplied time and resource budget, and return one result for every company. A timeout or missing result means that run is not scored; a Builderr infrastructure failure is rerun.

How we rank qualified entries: We average every scheduled daily batch while your submitted code is active. If your code fails or misses a batch after a clean evaluator rerun, that batch scores zero. A Builderr harness, infrastructure or shared-source failure is void and rerun with the same code. Ties go to fewer wrong-company facts, then better company coverage, then lower third-party API cost.

Full technical requirements and final ranking

Each official run has a fixed time and resource budget. Builderr supplies the company batch and output format for the run.

Return exactly one terminal result envelope for every input company. Use distinct available, not_available, blocked, not_applicable, ambiguous and failed states. Refresh must be idempotent: the same source snapshot must not create duplicate records or false changes.

External company recall and exact-company precision contribute to the score. A sparse batch is scored only on the evidence available for that batch; it does not create a new qualification threshold.

Exact tiebreak order: fewer wrong-company publications, then higher weighted company recall, then lower declared third-party cost. Document source rights, server-side secrets and safe URL handling. Supply pinned dependencies and one reproducible run command. Integrity, evidence and operational requirements in the evaluation contract still apply.

Your first submission is version 1. You may then submit up to four revised commit hashes before Oct 18, for five versions total. Each revision is frozen before its next official run and applies only to later batches. The evaluation contract contains the exact scoring terms and run format.

07 · Rewards

$2,500 in final prizes.

Main challenge — $2,000: $1,200 / $500 / $300 for first, second and third.

Separate JBOX bonus — $500: $250 / $150 / $100 for qualifying agents built with JBOX.

Community awards — $100 × 4: Sep 6, Sep 20, Oct 4 and Oct 18. Public voting never changes the technical ranking.

Only qualified entries enter the hosted Builderr gallery. Builderr covers hosting during the competition.

The winning builder gets the opportunity to partner with Håvard Liltved Dalen to launch Signalpost in Norway. Håvard is CPO at Fronted and co-founder of JBOX.

Build resources

Not ready to enter this one? Get challenges that fit your skills → Or join the Builderr Discord to ask questions or share what you are building → No answer there within a day? Email inquiries@builderr.ai.