AnyLearn
All interview prep
ProductMid-levelProduct Manager

Product Manager Interview Prep: Questions and a Mock Test

Product management interviews are the least standardised of any senior role, because the job differs so much between a platform team and a consumer app. What is stable is the assessment: interviewers are trying to work out whether you make decisions from evidence or from confidence. Almost every question in a product loop is a disguised version of "how do you know". This page covers the rounds, the answers that separate strong candidates, and ends with a graded mock across six areas.

The loop

How the process is structured

The interview loop: each round, how long it runs, and what it tests
RoundLengthWhat it tests
1.Product sense and design[2]Not publishedDesigning or improving a product under an open brief. Assessed on structure: goal, segment, problem, prioritisation, solution, measurement, in that order. Proposing features before naming a user and a problem is the most visible failure mode.
2.Analytical and metrics[3]Not publishedDefining success for a feature, explaining why a metric moved, or estimating a market. Also experiment interpretation, where the expectation is knowing what a p-value and a confidence interval do and do not license.
3.Execution and prioritisation[2]Not publishedTradeoffs under constraint: a slipping deadline, a large customer's request, technical debt against new features. The assessment is whether you can defend the call and say who owns the decision, not which framework you name.
4.Behavioural and leadership[1]Not publishedStructured stories about influence, disagreement and failure. Amazon's published principles are the clearest example of a written rubric, including Dive Deep, which expects leaders to be "skeptical when metrics and anecdote differ".

Bracketed markers point to the dated sources at the end of this article. Loops change; check the retrieval dates before relying on a round count.

Product sense is a structure, not an opinion

The design round, usually phrased as "design a product for X" or "how would you improve Y", is where most candidates lose the interview in the first two minutes by proposing features.

The structure interviewers are trained to listen for runs: clarify the goal, choose a user segment and say why, identify that segment's specific pain points, prioritise among them with a stated reason, then propose solutions to the top one, and only then say how you would measure whether it worked. Solutions arriving before a user and a problem is the single most common failure, and it is visible immediately.

Segmentation is where good answers separate. "All users" is not a segment, and choosing one narrow enough to have a distinct problem, then saying explicitly why you chose it over the alternatives, demonstrates the judgement being assessed. It is entirely acceptable to say that you are picking the smaller segment because the problem is more acute there and the learning transfers.

Breadth then depth is the pacing that works. Generate several distinct solution directions, not variations of one, then commit to one and go deep on how it works and what could go wrong. Candidates who generate ten shallow ideas read as unable to decide; candidates who commit instantly read as unable to consider alternatives.

Finally, say what would make you wrong. Naming the assumption your idea depends on is the difference between a proposal and a plan.

Discovery, and the outcome-versus-output distinction

Marty Cagan's 2014 article Good Product Team, Bad Product Team is quoted often enough in product interviews that its framing is effectively part of the vocabulary, and the contrasts are sharp enough to be worth knowing.

On what teams celebrate: "Good teams celebrate when they achieve a significant impact to the business (outcome). Bad teams celebrate when they finally release something (output)." That distinction underlies a lot of interview scoring. A candidate whose accomplishment stories are lists of shipped features is describing output; a candidate who says what changed for users and for the business is describing outcome.

On planning: "Good teams are skilled in the many techniques to rapidly try out product ideas to determine which ones are truly worth building. Bad teams hold meetings to generate prioritized roadmaps." This is not an argument against roadmaps so much as against treating a prioritised list as evidence.

On the working relationship: "Good teams have product, design and engineering sit side-by-side, and embrace the give and take between the functionality, the user experience and the enabling technology", and they "engage directly with end-users and customers every week".

The practical interview consequence is that discovery questions want you to name the riskiest assumption and the cheapest way to test it. A five-user interview round, a fake door test, a concierge version done by hand, or a prototype are all legitimate, and choosing among them by what risk they retire is the answer that scores.

Prioritisation you can actually defend

Prioritisation questions are rarely about the framework. Any of the common ones will do, and reciting one without applying it is a weak answer. What interviewers push on is the input you cannot know precisely.

The strongest structure is to make the comparison explicit: what is the expected impact, on which metric, for how many users, and what is the confidence in that estimate. Scoring frameworks that include a confidence term are useful precisely because they force you to admit that one of your numbers is a guess. Effort belongs in the comparison too, and getting it from engineering rather than inventing it is part of the job.

Expect a scenario where the answer is not to build anything. A request from the largest customer that would serve only that customer, a feature that would raise a metric while damaging trust, an item whose real cost is the maintenance nobody priced. Being willing to say no, and to say it with a reason the requester can understand, is explicitly assessed in most loops.

Technical debt and platform work are a common probe. The weak answer allocates a fixed percentage of capacity because it sounds balanced. The stronger answer treats debt as a set of specific items with specific costs, some of which are worth paying down now because they are slowing delivery measurably, and some of which are fine to leave.

And expect the deadline question. When scope, time and quality cannot all hold, saying which one you would move and who owns that decision is the point.

Metrics, and how they get gamed

The analytical round asks you to define success or to explain a movement, and the traps are consistent.

Defining success starts with the decision the metric informs. A metric nobody would act on differently at different values is a report. From there, name a primary metric, guardrails that catch the harm the primary is blind to, and a counter-metric where the obvious way to move the primary is undesirable. Engagement metrics are the classic case: notification opens rise when you send more notifications, so the guardrail is opt-out and uninstall rate.

Be precise about the shape of the metric. A ratio can move because its numerator moved, because its denominator moved, or because the mix underneath shifted while every segment stayed flat. That last case reverses conclusions and is worth naming explicitly.

On experiments, the expectation is that you know what a result does and does not license. A p-value is computed assuming the null is true, so it is not the probability the feature works. Checking daily and stopping at significance inflates false positives. Testing twenty metrics at a five percent threshold produces one significant result by chance. And a non-significant result is not evidence of no effect, particularly at small sample sizes.

The practical answer interviewers like: report the confidence interval and the decision's cost structure, not the verdict. Whether to ship a change whose interval spans a small harm and a meaningful gain depends on how reversible the change is.

Behavioural rounds, and what a good story contains

The behavioural round is not a formality. In many companies it carries as much weight as the product rounds, and Amazon is the clearest published example of a company that has written down what it assesses.

Amazon publishes sixteen Leadership Principles, and several map directly onto product judgement. Customer Obsession states that "leaders start with the customer and work backwards. They work vigorously to earn and keep customer trust." Dive Deep says leaders "stay connected to the details, audit frequently, and are skeptical when metrics and anecdote differ", which is a direct instruction to interrogate a number that disagrees with what users say. Are Right, A Lot asks that leaders "seek diverse perspectives and work to disconfirm their beliefs". And Bias for Action puts the reversibility argument plainly: "Speed matters in business. Many decisions and actions are reversible and do not need extensive study. We value calculated risk taking."

Whichever company you are interviewing with, the story structure that works is the same: situation, what you specifically did, and a quantified result. The most common defect is the first person plural. "We decided" tells an interviewer nothing about you; "I argued for X against Y because Z" does.

Prepare stories about disagreement and about failure, because both are asked in almost every loop, and the failure story should contain a real failure with a cost, not a disguised strength.

Open-ended

What they actually ask

  1. 1.How would you improve the experience of booking a train ticket?

    What a strong answer covers

    Strong answers clarify the goal first: is this about conversion, about support cost, about trust after a bad experience. Then a segment chosen deliberately, commuters buying the same journey repeatedly, occasional leisure travellers, and people travelling with constraints such as accessibility or luggage, with a stated reason for the choice. Then that segment's specific pains, which for repeat commuters is usually the number of steps to repeat a known purchase, and for occasional travellers is usually uncertainty about which ticket is valid and whether a cheaper option was missed. Only then solutions, several distinct directions rather than variations, one committed to and developed. The close is measurement plus the assumption that would make the idea wrong.

  2. 2.Your engineering lead says the feature will take three months, not the three weeks you planned. What do you do?

    What a strong answer covers

    The expected first move is to understand the estimate rather than negotiate it, because a tenfold gap usually means the requirement was understood differently rather than that anyone is being difficult. Then decomposition: what is the smallest version that tests the riskiest assumption, and what is driving the cost, often a non-functional requirement nobody stated such as migrating existing data or supporting an older client. Strong answers present options with consequences to whoever owns the deadline rather than choosing silently, and are explicit that scope is the variable they control while quality is not. The clearest signal of seniority is treating the engineer as the source of truth on effort and themselves as the source of truth on what can be cut.

  3. 3.Define success for a feature that lets users save articles to read later.

    What a strong answer covers

    Good answers avoid measuring the feature's own usage, since saves can be increased by prompting more aggressively while nobody reads anything. The outcome is articles actually read from the saved list, and beyond that whether saving is associated with retention. Strong candidates name a primary metric, guardrails such as complaint or opt-out rate, and a counter-metric for the obvious perverse incentive. They distinguish leading indicators available within days from lagging ones like retention that take weeks, and they raise novelty effects, since a new feature's usage spikes and decays. The best answers say what number would cause them to remove the feature, which is what makes the metric a decision rather than a report.

  4. 4.Your biggest customer demands a feature that only they would use. How do you respond?

    What a strong answer covers

    Interviewers are checking whether commercial pressure overrides product judgement. Strong answers separate the request from the underlying need, since customers describe solutions and the need beneath it is often shared. They quantify: revenue at risk, the cost to build, and the ongoing maintenance cost, which is the one most often omitted. They consider alternatives that serve the need without a bespoke build, including configuration, an integration or professional services. Where the answer is no, they say it with a reason and an alternative rather than deflecting. A candidate who simply builds what the largest account asks for has described account management, not product management.

  5. 5.Tell me about a time you were wrong about a product decision.

    What a strong answer covers

    This has to be a real failure with a real cost, not a strength in disguise. The structure that works: what you believed and why it was reasonable at the time, what you did, what the evidence eventually showed, and what specifically changed in how you work. The most valuable detail is usually the signal that was available earlier and that you did not weight properly, because it demonstrates the ability to audit your own reasoning rather than just the outcome. Answers that blame stakeholders or a shifting market read poorly. Answers that end with a general lesson such as talk to users more read as unearned; a specific change, such as always running the assumption test before committing engineering time, reads as real.

  6. 6.An A/B test shows your new onboarding raises signups by 12 percent but reduces 30-day retention by 4 percent. What do you do?

    What a strong answer covers

    The expected recognition is that this is a guardrail firing, and that the primary metric alone would have shipped a change that makes the business worse. Strong answers convert both into the same unit before deciding: 12 percent more signups retaining at a lower rate may still be more retained users, or may not, and the arithmetic settles it rather than intuition. They then investigate mechanism, usually that easier signup admits less committed users, which would be visible as a change in the composition of new cohorts rather than a change in behaviour within them. That distinction matters, because if the retained-user count rises and the rate falls only through mix, the change may be genuinely good and the metric misleading. Naming that possibility explicitly is the strongest available answer.

Worked examples

Three sample questions, answered

These three show the level the mock is pitched at, with the answer and the reasoning in the open. The graded paper keeps its answer key server-side.

1.In a product design round, what should come before proposing any solution?
Product sense and framing
  • An estimate of engineering effort
  • A competitive analysis of similar products
  • A clarified goal, a chosen user segment, and that segment's specific problem
  • A proposed success metric with a target value

Why: Solutions are only assessable against a problem, so an idea offered before a user and a goal cannot be evaluated by either party. Effort, competitors and metrics all have their place, but each of them presupposes that you have decided who you are building for and what you are solving.

2.According to Cagan's contrast, what do good teams celebrate?
Discovery and customer research
  • Shipping on the date they committed to
  • Achieving a significant impact on the business, which is an outcome
  • Completing every item on the quarterly roadmap
  • Reducing the size of the backlog

Why: The stated contrast is that "good teams celebrate when they achieve a significant impact to the business (outcome). Bad teams celebrate when they finally release something (output)." This is why accomplishment stories built from shipped features score worse than stories built from what changed for users.

3.Notification opens are your primary engagement metric. What is the most important guardrail?
Metrics, launch and iteration
  • Opt-out and uninstall rate
  • Average time to open a notification
  • Number of notifications sent per user
  • Click-through rate on notification content

Why: The obvious way to raise opens is to send more notifications, which harms users while the primary metric improves. The guardrail must capture that harm, and opt-outs and uninstalls are the direct measure of users voting against the tactic. Send volume is the input being gamed, not the harm.

The mock

An 18-question knowledge check

This is a knowledge check, not a simulation. The real loop happens on a whiteboard, in an editor, and in conversation. What this paper does measure is the underlying knowledge those rounds draw on: each question is tagged with a topic, grading happens per topic, and a weak topic points you at the course that fixes it.

Your paper0 / 18 answered
  1. 1.A candidate answers 'design a fitness app' by choosing 'all users who want to get fit'. What is the weakness?
    Product sense and framing
  2. 2.Why is it a strong move to state the assumption your proposed solution depends on?
    Product sense and framing
  3. 3.You have three distinct solution directions and limited time. What is the strongest way to proceed?
    Product sense and framing
  4. 4.Which discovery method best tests whether users want a feature before it is built?
    Discovery and customer research
  5. 5.What does a concierge test involve?
    Discovery and customer research
  6. 6.Cagan contrasts good and bad teams on planning. What is the stated contrast?
    Discovery and customer research
  7. 7.Why do scoring frameworks that include a confidence term tend to produce better decisions?
    Prioritisation and roadmaps
  8. 8.What is the strongest way to handle a request for technical debt work in a prioritisation discussion?
    Prioritisation and roadmaps
  9. 9.A stakeholder insists a feature ships by a fixed date and the scope will not fit. Which variable is normally the right one to move?
    Prioritisation and roadmaps
  10. 10.Conversion rate rose while total conversions fell. What is the most likely explanation?
    Metrics, launch and iteration
  11. 11.What is a counter-metric, as distinct from a guardrail?
    Metrics, launch and iteration
  12. 12.Which statement about a metric definition is the strongest test of whether it is worth tracking?
    Metrics, launch and iteration
  13. 13.An experiment returns p = 0.02 on the primary metric. Which conclusion is warranted?
    Experiments and evidence
  14. 14.Why is stopping an experiment as soon as it reaches significance a problem?
    Experiments and evidence
  15. 15.Engagement with a new feature spikes for two weeks then returns to baseline. What is the most likely cause?
    Experiments and evidence
  16. 16.In a behavioural answer, why is 'we decided to rebuild the checkout' a weak formulation?
    Stakeholders and influence
  17. 17.Amazon's Dive Deep principle states that leaders are 'skeptical when metrics and anecdote differ'. What behaviour does that imply?
    Stakeholders and influence
  18. 18.You have no authority over the engineering team but need a change prioritised. What is the strongest approach?
    Stakeholders and influence
18 questions left to answer.
Apparatus

Sources

Hiring loops change. Every claim above carries a retrieval date so you can judge how current it is.

  1. [1]Amazon, Leadership Principles · retrieved 2026-08-13
  2. [2]Marty Cagan, Good Product Team / Bad Product Team, Silicon Valley Product Group, June 2014 · retrieved 2026-08-13
  3. [3]Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments, and the experimentation paper archive · retrieved 2026-08-13
Keep preparing

Refresh your memory

Free learning paths covering the ground this loop tests, whatever your score. Each one ends with a shareable certificate.