AI & Technical question

A frontier lab says its model scores well on public benchmarks but still fails on long-horizon, repo-scale engineering tasks. How would you design a contamination-resistant evaluation or RL environment that surfaces those failures, while preserving reproducibility, trustworthy reward signals, and automated code-correctness verification?

Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.

Start a mock interview on this question · Mock interview from a job description

The answer guide for this question is on its way. You can still practice it now.

More ai & technical questions

More questions from Scale AI

Learn the skill behind it

Chapters of the AI PM course that teach what this question tests.

Preparing for a specific role?

Book summaries for this kind of question

Browse all 4,000+ questions in the bank