OpenRefinery · Founding Benchmark Fellow
A founding cohort of graduate researchers, each building the benchmark for their science.
Papers pay $100–$500 by difficulty and quality, alongside your program. Remote, no exclusivity, work on your own schedule.
Frontier labs already pay to climb benchmarks in coding and math. The sciences are next, and the question is who writes them. We think the answer is a researcher at candidacy depth who uses these models on real work every day, not a contractor following someone else's rubric. So we're hiring Fellows, field by field, to build their field's benchmark and put their name on it.
This is not annotation work. We built annotation at scale at Turing, so we know exactly what that assembly line looks like, and this role exists because that model can't produce what labs need next. Nobody hands you a rubric here; you write it. If a task ever feels like labeling, the role design has failed and we want to hear about it.
What a Fellow owns
The map of your field: the problem areas that matter most (we think it's about five; you'll tell us) and the taxonomy every research trace gets tagged against. The standards for turning raw traces into benchmark samples labs score against. When your field's benchmark ships, the judgment in it is yours.
Where the data comes from, before your advisor asks
Traces come from our researcher network: scientists who run our open-source capture tool on their own AI sessions, redact locally on their own machines, and approve every session before it uploads. Nothing enters the pipeline that its owner didn't explicitly approve. Being a Fellow doesn't require contributing your own research data.
Credit, concretely
Each field's benchmark ships as a public release with named contributor credit for the Fellow who built it. The commercial dataset behind it stays ours; the public artifact carries your name.
The level we're hiring at
You've pushed a real research problem to candidacy or beyond. You could referee a paper in your subfield and catch the error the authors missed. AI agents like Claude Code or Codex are already load-bearing in your research (for most people that's 10 or more hours a week), and you can tell genuinely correct from merely plausible without running the code. Senior PhD students are explicitly welcome; depth matters more than years.
Fields
Fellows in materials science, physics, chemistry, biology (including computational), applied math, EE, and CS/ML — no cap per field. If your field isn't listed and your research runs on AI agents, make the case.
How it works
- Remote, with no exclusivity. Built to coexist with your program.
- Hours are yours to set, no minimum and no cap. Disappear for a deadline week, or make this your main commitment outside research. Both work.
- Paid per accepted paper: $100–$500, by difficulty and quality.
- F1/visa-friendly payment structure available.
- You work directly with our founders and with researchers at frontier AI labs.
- If you know researchers who belong in the network, there's budget to bring them in. Recruiting is optional, not the job.
How the application works
- Apply (10 minutes). The question we read first: one thing models consistently get wrong in your area.
- Initial screen, within 48 hours of submitting. Partly automated (fixed criteria on your written answer, with a human-review path — see the privacy notice). Outcomes: move forward, waitlist with a stated review date, or a clear no.
- Verification. We confirm your institution and identity (institutional email, ORCID, or a quick manual check).
- Work sample. A short paper-reproduction task, graded against the paper's own numbers. No deadline — take it at your own pace. Passing earns a $50 credit, paid with your first paper payout after you sign on.
- Agreement and first paid paper. Click-wrap contributor agreement, payout profile, and you're in the paid queue: $100–$500 per accepted paper.
Who's behind this
The three of us built Turing's data business together and scaled the company to a $2.2B valuation and roughly $300M ARR. Kai built the human data annotation platform used by Facebook AI Research (FAIR) and led the data team at Zoox. Char was Turing's second employee, led product from pre-seed through the Series E, and has exited two companies. Alex ran Turing's APAC data business.
Backed by Audacious Ventures (investors in Reflection AI and Decagon) and Zetta Venture Partners (early investor in Kaggle, acquired by Google, and in Domino Data Lab; Zetta's founder chairs the MIT Corporation, MIT's board of trustees), plus AI researchers from OpenAI, Anthropic, Nvidia, and Meta. Our capture tool runs with researchers at Stanford, MIT, Caltech, ETH Zurich, UPenn, Berkeley, NUS, and more.
Questions: contact@openrefinery.ai