AI & Technical question

Glean Model Hub may support multiple LLMs for search, assistant, and agent workflows. Design an evaluation framework you would use to decide whether a new model should be added to the portfolio, including how you would assess answer quality, latency, cost, safety, and enterprise-specific requirements such as permissions, grounding, and consistency.

Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.

Start a mock interview on this question · Mock interview from a job description

What this question tests

Tests ability to design a rigorous, enterprise aware model evaluation framework covering quality, cost, safety, and Glean specific requirements.

How to approach it

  1. Define what the evaluation decides: whether a candidate model earns a slot in Model Hub for search, assistant, or agent workflows.
  2. Build an offline eval suite with a golden set of real enterprise queries, scored on accuracy, groundedness to cited sources, and latency under load.
  3. Add enterprise checks: does the model respect document level permissions, does it stay consistent in tone and format across repeated queries.
  4. Layer in safety testing, for example red teaming for prompt injection through retrieved documents and cross workspace data leakage.
  5. Run a limited live pilot with a design partner workspace comparing the candidate against the current default on real usage, not just offline scores.

What a strong answer includes

Common mistakes

Likely follow-up questions

More ai & technical questions

More questions from Glean

Learn the skill behind it

Chapters of the AI PM course that teach what this question tests.

Preparing for a specific role?

Book summaries for this kind of question

Browse all 4,000+ questions in the bank