How would you test a vending machine — and anything else they ask
"How would you test a pen?" isn't about the pen. It's whether you ask what kind, for whom, before you start listing. Here's the method, then fifteen objects worked through with it.
There's no right answer, so it can't be memorised — which is exactly why interviewers use it. Almost everyone fails the same way: they start listing. "Check it dispenses the item, check the buttons…" Four minutes later the interviewer still can't tell whether you covered the space or ran out of ideas.
The method: clarify, cover, close
Three moves that work for a vending machine, a chair or a search box. The seven axes are the part worth memorising.
- 1ClarifyAsk 3–4 questions before a single test case. This is the step being scored.
- 2Cover the axesSeven directions that apply to any object, physical or digital.
- 3CloseState your assumptions and what you'd test first if time were short.
Say the axes out loud as headings — "let me take function, then limits, then misuse…" — then stop and ask which one they want in depth. That question turns a monologue into a conversation.
Machines that take money
How would you test a vending machine?
Clarify what it sells, what it accepts and where it stands — then walk seven axes: function, limits, abuse, failure, stress, human factors and environment. The money and the stock are what make it interesting: exact change, a jam mid-dispense, and a power cut after payment.
- What does it sell — snacks, drinks, tickets? Perishable or not?
- What does it accept — coins, notes, card, phone, staff badge?
- Where does it stand — indoor office, station platform, outdoors?
- Is it connected — does it report stock and take remote price updates?
- Function
- Select a product, pay, get exactly that product, get correct change, get a receipt if it offers one. Then the same for every payment type it accepts.
- Limits
- Exact money, one unit short, one over. Last item in a slot. Empty slot. Full change hopper, and an empty one — does it refuse the sale or take money it can't give change for?
- Abuse
- Foreign coins, torn notes, a note pulled back on a string. Two selections at once. Pressing cancel exactly as it dispenses. Tilting or shaking it for a second item.
- Failure
- Power cut after payment, before dispense — does it refund or remember? Product jams halfway. Card reader online, network dead. Coin path blocked. Every case asks: who ends up out of pocket?
- Stress
- Continuous purchases until a slot empties. A full restock cycle. Cash box near capacity. All motors run for hours in a queue at a station.
- Human
- Readable at a glance, reachable from a wheelchair, buttons usable with gloves, clear error text, sensible feedback when it takes your money and thinks.
- World
- Cold platform vs warm lobby (condensation on the glass, chilled drinks). New coin design. Price change day. Restock access without exposing cash. Audit log for the operator.
Power fails in the second between payment taken and product dropped. A candidate who names this — and asks whether the machine refunds, remembers, or silently keeps the money — has just described the only failure that reaches the news.
The trapListing 30 button-press cases and never mentioning money, stock or power. The vending machine is a payment system that happens to hold crisps.
How would you test an ATM?
Same seven axes, higher stakes: every failure path ends in someone's money. Clarify the card types and limits, then push hardest on Failure — cash dispensed but not debited, debited but not dispensed, card retained, session timeout mid-transaction.
- Which cards — own bank, network partners, international?
- What can it do — withdraw, deposit, transfer, PIN change, statements?
- What are the limits — per transaction, per day, note denominations?
- Is it in a branch lobby or on an open street?
- Function
- Card in, PIN, withdraw, cash out, balance correct, card returned, receipt correct. Then every other operation it offers, each verified against the account afterwards — not just the screen.
- Limits
- Minimum withdrawal, maximum, daily cap, one over the cap. Amount not divisible by available notes. Balance exactly equal to the request. Zero balance. Account at overdraft edge.
- Abuse
- Wrong PIN three times — retain or return? Damaged or cloned card. Cancel at every step. Walk away mid-session. Two cards in quick succession. Someone shoulder-surfing: does the screen mask?
- Failure
- Power or network loss after debit, before dispense. Cash jam. Printer out of paper — does the transaction still complete? Cash cassette empty mid-count. Every one asks: does the ledger match the cash drawer afterwards?
- Stress
- Payday queue, back-to-back sessions for hours. Cassettes down to the last notes. Reconciliation after a thousand transactions.
- Human
- Timeout long enough to read but short enough to be safe. Audio jack and screen-reader mode. Reachable keypad. Language choice before PIN. Clear cancel at every step.
- World
- Outdoor cold and rain on the card slot. Currency and note designs. Regulatory audit trail. Camera and lighting requirements. Servicing without opening the safe.
Debited but not dispensed. Say the sentence 'the ledger and the cash drawer must reconcile after every failure path' and you have answered this question better than most candidates manage in five minutes.
The trapTesting the screens and forgetting the account. Every ATM case has two results — what the user sees and what the bank records — and only one of them is on the display.
How would you test a coffee machine?
Clarify whether it's an office bean-to-cup or a paid commercial unit — that decides whether money is in scope. Then function, limits (empty water, no beans, full grounds), abuse (cup removed mid-pour), failure (power cut mid-brew) and hygiene, which is this machine's unique axis.
- Free office machine or paid? Payment changes half the test plan.
- Bean-to-cup, pods, or instant? Different consumables, different failures.
- Does it plumb into mains water or use a refillable tank?
- Who cleans it, and how often — is descaling in scope?
- Function
- Every drink on the menu, correct volume, correct temperature, correct strength. Milk options. Cup size options. Sugar. Then the same drink twice in a row — consistency is the product.
- Limits
- Water tank empty, and one cup's worth left. Beans out. Grounds tray full. Milk empty mid-pour. Largest cup at maximum strength. Smallest cup.
- Abuse
- Cup removed mid-pour. No cup at all. Wrong-size cup under the spout. Two drinks selected at once. Cancel mid-brew. Pod inserted upside down or already used.
- Failure
- Power cut mid-brew — does it clear the group head or leave hot slurry? Blocked spout. Pump failure. Door opened while hot. Does it stop safely and say why?
- Stress
- Twenty cups in the 9am queue. Continuous brewing until thermal cut-out. Two weeks without cleaning. Scale build-up in hard water.
- Human
- Buttons legible while half-awake, hot surfaces marked, drip tray removable without spilling, error messages that say what to refill rather than a code.
- World
- Hygiene and cleaning cycles — the axis unique to this machine. Milk left in the line overnight. Descaling reminders. Waste and recycling of pods. Noise in an open-plan office.
Milk. It is the only ingredient that spoils, blocks lines, and becomes a health issue if the cleaning cycle is skipped. Mentioning the milk line is the tell of someone who has thought past the buttons.
The trapTreating it as a drinks menu. Hygiene, consumables and thermal safety are where a real coffee machine's defects live.
Everyday objects
How would you test a pen?
The classic. It is not about the pen — it is about whether you ask what kind of pen and who uses it before you start. Clarify first, then run the axes: writes on paper (function), ink runs out (limits), used as a lever (abuse), dropped (failure), left uncapped (world).
- What kind — ballpoint, gel, fountain, marker? Disposable or refillable?
- Who uses it, and for what — school, office, hospital charts, signing contracts?
- What is it meant to write on — paper only, or glossy, glass, skin?
- What is the price point? A promotional giveaway and a luxury pen have different acceptance criteria.
- Function
- Writes on the first stroke, writes continuously, produces a consistent line, colour matches the label, cap fits and clicks, clip holds to a pocket.
- Limits
- First stroke after months in a drawer. Last of the ink — does it fade or stop cleanly? Continuous line until empty. Maximum writing pressure. Shortest possible tap.
- Abuse
- Writing on the wrong surfaces. Pressing hard enough to tear paper. Chewed cap. Used to open a package. Dropped nib-down from desk height. Left in a car in summer.
- Failure
- Dropped and cracked — does it leak? Cap left off for a day. Stored nib-up vs nib-down. Air travel pressure change. Does failure make a mess or just stop?
- Stress
- Continuous writing for hours. Ten thousand cap clicks. Repeated drop cycles. Sits in a pocket for a year.
- Human
- Grip comfort over a long session, left-handed smudging, safe cap design for children, legibility of the ink for the person reading it.
- World
- Cold, heat, humidity. Ink permanence on a legal document. Safety standards for cap airflow. Refill availability. Recyclable body.
The vented cap. Pen caps have a hole in them because children swallow them — naming a safety standard turns a joke question into a demonstration that you think about the user, not the product.
The trapJumping straight into 'check if it writes'. The whole point of the pen question is whether you ask what pen and for whom before generating a single case.
How would you test a chair?
Clarify who sits in it and where — an office chair, a stacking chair and a child's chair have different pass criteria. Then load limits, stability under abuse, endurance over years, and safety, which is the axis this object is really about.
- What kind — office, dining, stacking, folding, child's, wheelchair?
- Who uses it, and for how long at a stretch?
- Where does it live — carpet, tile, outdoors?
- Is it adjustable, and does it have wheels or arms?
- Function
- Takes a seated adult, stays level, height and tilt adjust and hold, wheels roll and lock, arms take weight, it stacks or folds if it claims to.
- Limits
- Stated maximum load, and just over it. Minimum user size. Lowest and highest seat settings. Maximum recline. Fully folded and fully extended.
- Abuse
- Leaning back on two legs. Standing on it. Sitting on the arm. Dragging it by one arm. Sitting down hard, repeatedly. A user well over the rated weight.
- Failure
- One caster removed, one bolt loose, gas lift blown. Does it degrade safely or collapse? Fabric torn — do sharp parts become exposed?
- Stress
- Eight hours a day for five years. Ten thousand sit-stand cycles. Rolling over a cable a thousand times. Stacked twenty high in storage.
- Human
- Lumbar support, edge pressure behind the knees, adjustment reachable while seated, no pinch points, no sharp edges at child height.
- World
- Carpet vs hard floor. Humidity and outdoor use. Flammability standards for office furniture. Weight for shipping. Assembly instructions and the parts in the box.
Leaning back on two legs is not misuse you can design out — everyone does it. A chair that fails there fails in the real world, which is why it belongs in the plan rather than in a 'user error' note.
The trapDescribing comfort as the main criterion. Comfort matters, but load, stability and safety are what make a chair pass or get recalled.
How would you test a calculator?
Clarify basic or scientific, physical or app — then this is the object where Limits and Abuse do the heavy lifting: divide by zero, overflow, floating-point rounding, and operator keys pressed in impossible orders.
- Basic four-function, scientific, or financial?
- Physical device or software? A phone app inherits the OS's problems.
- What precision is promised — how many digits, how does it round?
- Does it have memory, history, or a paper tape?
- Function
- Each operation alone, then chained. Order of operations. Memory store and recall. Percent, sign change, square root. Clear vs clear-entry — they are different, and people confuse them.
- Limits
- Largest and smallest representable number. Maximum digits on screen. Divide by zero. Zero divided by zero. Square root of a negative. Overflow and underflow — does it say so or show nonsense?
- Abuse
- Two operators in a row. Equals pressed first. Multiple decimal points. Holding a key down. Pressing two keys at once. A very long chain of operations.
- Failure
- Battery dying mid-calculation. App backgrounded and restored — is the entry preserved? Screen cracked but responsive. Does it lose memory silently?
- Stress
- Thousands of key presses. Repeated equals. Long-running chained operations. Memory used continuously.
- Human
- Key size and spacing, tactile feedback, screen legibility at an angle, screen-reader announcement of the result, undo for a mis-key.
- World
- Locale decimal separator — comma or point. Exam-mode restrictions. Battery or solar under low light. Regional standards for financial rounding.
0.1 + 0.2. If the device claims decimal precision and shows 0.30000000000000004, that is a real defect for a calculator even though it is correct floating-point arithmetic. Knowing why that happens is a genuine differentiator.
The trapTesting only arithmetic that works. The interesting cases are the ones with no clean answer — divide by zero, overflow, rounding — because they force a decision about what the product should say.
How would you test a toaster?
A small object with a real safety axis. Clarify the slots and settings, then function and limits are quick — the interesting work is Abuse and Failure, because the failure modes involve heat, electricity and things people push into slots.
- How many slots, and are they wide enough for bagels and thick bread?
- What settings — browning levels, defrost, reheat, cancel?
- Is it a pop-up or a conveyor type for commercial use?
- Is there a removable crumb tray?
- Function
- Each browning level produces a visibly different result, both slots behave the same, the lever latches, toast pops at the end, cancel stops immediately, defrost and reheat do what they claim.
- Limits
- Thinnest and thickest slice. Nothing in the slot. Setting 1 and maximum. Cancel one second in. Two slices of different thickness at once.
- Abuse
- Knife or fork inserted — is the element reachable? Oversized item jammed. Lever held down. Something with butter or cheese on it. Repeated immediate restarts.
- Failure
- Power cut mid-cycle — does the lever release or stay latched with bread inside? Element failing on one side. Thermostat failure. Crumb tray full and igniting. Each asks: does it stop before it becomes a fire?
- Stress
- Continuous use in a hotel breakfast service. Hundreds of cycles a day. Crumb build-up over months. Cord flex over years.
- Human
- Cool-touch exterior, lever within reach, settings legible, no need to reach into a hot slot, cord length and placement, auto-eject at the end.
- World
- Voltage variation, electrical safety certification, humidity in a kitchen, cleaning without immersing it in water, thermal cut-out standards.
The metal knife. Every tester jokes about it; the real question is whether the element is reachable and whether the appliance is designed so that reaching in is safe when it is switched off but still hot.
The trapProducing a list of browning levels. The safety axis — heat, electricity, crumbs and human fingers — is where a toaster's defects actually matter.
Machines with a motor
How would you test a washing machine?
Clarify the model and the programmes, then treat it as three systems that must agree: water, drum and heat. The best cases live in Failure — door opened mid-cycle, power cut with a full drum, drain blocked — because each one asks whether it fails safe with water and heat involved.
- Which programmes, and is there a dryer built in?
- Front or top loading? Door interlock behaves differently.
- Is it connected — app control, remote start, firmware updates?
- What is the rated load, in kilograms?
- Function
- Each programme end to end: fill, wash, rinse, spin, drain, finish signal. Correct temperature, correct duration, correct spin speed. Detergent drawer drawn at the right moment.
- Limits
- Empty drum. Single sock. Rated maximum load, and over it. Coldest and hottest programme. Longest and shortest cycle. Maximum spin with an unbalanced load.
- Abuse
- Door forced mid-cycle. Programme changed while running. Power switched off and straight back on. Overdosed detergent. Coins and keys in the drum. Child pressing every button.
- Failure
- Power cut with a full drum of water — where does the water go on restart? Drain blocked. Water supply closed mid-fill. Heater fails. Unbalanced load at 1400rpm. Each asks: does it stop safely and drain, or flood the kitchen?
- Stress
- Back-to-back cycles for days. Maximum load every time. Ten years of door-seal cycles. Continuous vibration on an uneven floor.
- Human
- Door interlock so a child cannot open it mid-cycle, controls readable without the manual, end-of-cycle signal audible from another room, error codes that mean something.
- World
- Hard water and scale. Voltage variation. Cold inlet only vs hot and cold. Energy-label compliance. Vibration transmitted through an apartment floor. Servicing access to the pump filter.
Power cut with a full drum. On restart, the machine must know there is water inside and drain or resume — not open the door. That single case is where safety, state persistence and hardware all meet.
The trapListing the programmes and stopping. Water plus heat plus a locked door means the failure paths are the test plan, not an appendix to it.
How would you test a microwave oven?
Clarify the power levels and features, then note the one axis that outranks everything: safety interlocks. Door opened while running must cut the magnetron instantly — everything else is secondary to that, and saying so first is the strongest possible opening.
- What power levels and preset programmes?
- Does it grill or convect as well as microwave?
- Is there a turntable, a weight sensor, an inverter?
- Domestic or commercial — a staff canteen unit gets abused far harder.
- Function
- Set time and power, it heats for exactly that long at that level, turntable rotates, light and fan run, it beeps and stops. Every preset. Defrost by weight.
- Limits
- One second. Maximum time. Zero seconds — does it refuse? Power level 1 and level 10. Empty cavity. Maximum rated load. Timer rolling past its maximum digits.
- Abuse
- Metal inside. Door opened repeatedly mid-cycle. Start pressed with the door open. Time changed while running. Running empty. Buttons mashed. Something that boils over.
- Failure
- Power cut mid-cycle — does it resume or reset? Door switch failing. Turntable motor jammed. Magnetron overheating. Every case must end with 'the emission stops'.
- Stress
- Continuous use through a lunch rush. Maximum power for the maximum duration, repeatedly. Ten thousand door cycles. Grease build-up over a year.
- Human
- Door opens easily but latches securely, hot surfaces marked, controls usable by an older user, clear feedback that it has finished, child lock.
- World
- Radiation leakage limits — a regulated, measurable requirement. Voltage variation. Ventilation clearance around the unit. Cleaning without damaging the waveguide cover.
Open the door one millisecond into the cycle and the magnetron must already be off — the interlock is tested before anything else, because it is the only failure here that injures someone.
The trapTreating it as a timer with a light. A microwave is a regulated emitter behind a safety interlock; that framing is what separates a QA answer from a user-manual answer.
How would you test a lift (elevator)?
Clarify the building, the floors and the capacity, then lead with safety and failure — overload, door obstruction, power loss between floors, emergency call. A lift is a safety system that also moves people between floors, and interviewers listen for which of those you name first.
- How many floors, how many lifts, and is there a basement or roof access?
- Rated capacity in people and kilograms?
- Any special modes — fire service, goods, card-restricted floors?
- What happens on power loss — battery lowering, generator, none?
- Function
- Call from every floor in both directions. Travel to every floor. Doors open and close fully. Correct floor indicator and direction arrows. Sensible queuing when several calls arrive at once.
- Limits
- Empty. Rated capacity exactly. One person over — does it refuse to move and say why? Top and bottom floor. All buttons pressed at once. Longest possible journey.
- Abuse
- Doors blocked repeatedly. Button held down. Jumping inside the car. Every floor selected then cancelled. Forcing doors open between floors. Pressing call on every floor simultaneously.
- Failure
- Power loss between floors — does it lower to the nearest floor and open? Door sensor failure. Overload sensor failure. Emergency call button with the network down. Each case must end with people getting out.
- Stress
- Morning rush in an office tower, continuous cycles. Maximum load every trip. Door cycles over years. Two lifts, one out of service.
- Human
- Braille and audible floor announcements, button height reachable from a wheelchair, door timing for a slow walker, mirror and handrail, clear emergency instructions.
- World
- Fire-service mode recall. Statutory inspection intervals. Temperature and humidity in the shaft. Machine-room access. Compliance certification — this is one of the few objects where the regulator writes some of your test cases.
Overload plus a closing door plus a call already queued. Real lift defects live in the interaction of subsystems, not in any single one — naming an interaction rather than a feature is what marks a senior answer.
The trapTesting it like a set of buttons. Safety interlocks, capacity and evacuation are the substance; floor selection is the easy part.
How would you test an escalator?
Similar to a lift but continuous, so Stress and Human become the leading axes: it runs all day and people step on and off it while it moves. Emergency stop, comb-plate obstruction and direction switching are the cases that matter.
- Indoor or outdoor, and what rise and speed?
- Does it reverse direction, or run one way only?
- Does it idle and start on approach, or run continuously?
- What is the expected footfall — shopping centre or metro station?
- Function
- Runs at rated speed in the intended direction, handrail moves in sync with the steps, lighting and direction indicators correct, sensor-start works if fitted.
- Limits
- One person. Full capacity, every step occupied. Heaviest permitted item on the steps. Slowest and fastest permitted speed setting.
- Abuse
- Running up the down escalator. Sitting on the handrail. A trolley or pushchair on the steps. Something dropped into the comb plate. Emergency stop pressed for fun during peak flow.
- Failure
- Power loss with people on it — does it coast or brake, and how hard? Handrail stops while steps continue. Step chain breaks. Comb plate obstructed. Each asks: does it stop in a way that does not throw people forward?
- Stress
- Sixteen hours a day, seven days a week. Peak footfall for an hour. Years of step-chain cycles. Weather and grit if it is outdoors.
- Human
- Handrail synchronisation is a genuine safety issue, step edges visible, entry and exit plates safe for heels and pushchairs, audible and visible warnings, emergency stop reachable but not accidental.
- World
- Outdoor rain and ice. Grit wearing the steps. Statutory inspection. Cleaning while out of service. Noise. Access for maintenance without closing the whole concourse.
The handrail running slightly slower than the steps. It sounds trivial and it is a documented injury cause, because a hand travelling slower than feet pulls a person off balance. That is the kind of detail that ends the interview well.
The trapTreating it as a moving staircase with an on switch. The interesting behaviour is what happens when it stops with people on it.
How would you test a traffic light?
Clarify the junction and the modes, then the whole answer turns on one thing: no two conflicting directions may ever be green. Everything else — timings, pedestrian phases, sensor inputs — sits under that invariant.
- How many directions, and are there filter arrows or pedestrian crossings?
- Fixed timing, sensor-driven, or centrally controlled?
- Is there a night or low-traffic mode, and an emergency-vehicle override?
- What is the failure mode by design — all red, or flashing amber?
- Function
- Correct sequence per direction, correct durations, pedestrian phase with its beeper and countdown, filter arrows, sensor-triggered changes, night mode.
- Limits
- Minimum green time. Maximum wait before a direction gets a turn. Pedestrian button pressed once vs fifty times. Continuous traffic on one arm and none on the other.
- Abuse
- Every pedestrian button held down. A vehicle stopped on the sensor loop permanently. Emergency override triggered repeatedly. Conflicting central commands sent at once.
- Failure
- Lamp burnt out — does the controller detect it? Power cut and restart: what does it show while it re-syncs? Sensor stuck. Communication with the central system lost. Each must degrade to a documented safe state.
- Stress
- Rush hour on all arms. Continuous operation for months. Rapid mode switching. All pedestrian phases requested every cycle.
- Human
- Visible in low sun and heavy rain, audible signal for blind pedestrians, tactile cone under the button, crossing time long enough for a slow walker, countdown accuracy.
- World
- Regional signal sequences differ by country — red-amber before green exists in some, not others. Weather, vandalism, lamp lifespan, and the standards body that defines every timing you were about to invent.
The safety invariant: assert that no two conflicting directions are ever green, in every mode, during every transition, including power-up and fault. That single sentence is worth more than fifty timing cases.
The trapTesting colours and durations while never stating the invariant. The whole system exists to prevent one specific event, and naming it is the answer.
Software people actually ship
How would you test a search box?
Clarify what is being searched and how results rank, then the seven axes translate directly: exact and partial matches (function), empty and very long queries (limits), injection and wildcards (abuse), backend down (failure), and relevance — the axis that has no single right answer.
- What is being searched — products, documents, users? How large is the corpus?
- Is it exact match, fuzzy, or full-text with ranking?
- Are there filters, autocomplete, pagination, saved searches?
- What is the expected latency, and is relevance measured anywhere?
- Function
- Exact match returns the item. Partial and multi-word queries. Case insensitivity. Autocomplete suggestions. Filters combined with a query. Pagination consistent across pages. Result counts correct.
- Limits
- Empty query. One character. Maximum-length query. A query matching everything. A query matching nothing — is the empty state helpful? First and last page. Exactly one result.
- Abuse
- SQL and script payloads. Wildcards and regex characters. Unicode, emoji, right-to-left text. Leading and trailing spaces. A thousand-word paste. Rapid repeated submissions.
- Failure
- Search backend down — does the page degrade or hang? Timeout mid-query. Partial index after a rebuild. Stale cache serving deleted items. Network lost during autocomplete.
- Stress
- Concurrent searches at peak. Very large result sets. Index rebuild while serving. Autocomplete firing on every keystroke — is it debounced?
- Human
- Keyboard-only use, results announced to a screen reader, visible focus, obvious empty state, spelling suggestions, and latency that feels instant rather than merely fast.
- World
- Locale and language stemming. Accented characters matching unaccented queries. Permissions — search must never return a document the user cannot open. Analytics on zero-result queries.
Permission-aware results. A search that returns titles of documents the user is not allowed to read is a data leak that every functional test passes. It is the single most valuable case on this list.
The trapTesting that search finds things. The defects are in what it should not return, what it does with nothing, and what it does when the index is behind.
How would you test a shopping cart and checkout?
Clarify the payment methods and the inventory model first — those two decide most of the plan. Then the axes land hard on Failure and Abuse, because every path involves money and stock changing at the same time.
- What payment methods, and is there a saved-card or wallet flow?
- Is stock reserved when added to cart, or only at payment?
- Are there discounts, coupons, taxes, shipping rules, guest checkout?
- Single currency and region, or many?
- Function
- Add, update quantity, remove, empty cart. Totals with tax and shipping. Coupon applied and removed. Guest and logged-in checkout. Payment succeeds, order created, confirmation sent, stock decremented.
- Limits
- Empty cart checkout. One item. Maximum quantity. Stock of exactly one. Cart total of zero after a full discount. Maximum coupon value. Largest possible order.
- Abuse
- Coupon applied twice. Price or quantity tampered in the request. Back button after payment. Double-clicking pay. Two browser tabs buying the last item. Expired coupon reused.
- Failure
- Payment gateway timeout — charged or not? Network lost after payment, before confirmation. Stock sold out between cart and pay. Session expiring mid-checkout. Each asks: does the customer's money and the order state agree?
- Stress
- Flash sale on one product. Concurrent checkouts for the last unit. Cart with hundreds of lines. Payment provider slow but not down.
- Human
- Errors shown next to the field that caused them, progress visible, card form usable on mobile, screen-reader labels, no surprise costs at the final step.
- World
- Currency and tax rules by region. Address formats. Card types by country. Regulatory receipts. Timezone on order timestamps.
Two tabs, one unit of stock, pay in both. Whichever way the system resolves it must be deliberate — one succeeds and one fails cleanly with the money untaken. Concurrency on the last item is where real carts break.
The trapWalking the happy path and calling it done. Checkout is a distributed transaction across payment, stock and orders; the value is in the paths where those three disagree.
How would you test a messaging app like WhatsApp?
Clarify the scope — one-to-one, groups, media, calls — because the whole app is too broad for one answer. Then the axes: delivery states (function), group size and message length (limits), offline and reconnect (failure), and ordering, which is this product's hardest problem.
- Which part — one-to-one chat, groups, media, voice notes, calls, status?
- Which platforms, and does it sync across devices?
- Is it end-to-end encrypted? That constrains what a test can even observe.
- What are the limits — group size, file size, message length?
- Function
- Send, deliver, read receipts in the right order. Text, emoji, media, replies, forwards, deletes, edits. Group create, add, remove, leave. Notifications matching the app state.
- Limits
- Empty message. Maximum length. Maximum group size, and one over. Largest permitted attachment. Thousands of messages in one thread. Oldest message still loadable.
- Abuse
- Spam rate. Blocked user still trying. Malicious file types. Deleting a message the recipient already read. Editing after forwarding. Screenshot of a disappearing message.
- Failure
- Offline send — queued and delivered later, in order? Airplane mode mid-send. App killed while uploading media. Server unreachable. Recipient offline for a week. Message ordering after reconnect is the acid test.
- Stress
- Group of the maximum size all typing at once. Media-heavy thread. Years of history on a low-end phone. Poor network with constant reconnects.
- Human
- Readable at small sizes, screen-reader announcement of new messages, one-handed reachability, clear delivery states, notification that respects mute and do-not-disturb.
- World
- Timezones on timestamps, right-to-left languages, emoji rendering across platforms, storage limits, data-saver mode, regional feature and privacy rules.
Two people send while both are offline, then both reconnect. Message ordering under partition is the genuinely hard problem in any chat app, and naming it separates someone who has tested distributed systems from someone who has tested screens.
The trapTrying to test the whole app in one answer. Scope it out loud — 'let me take one-to-one messaging first' — and the interviewer will thank you for it.
Frequently asked questions
Why do interviewers ask how you would test a pen or a vending machine?
Because it has no right answer, so it can't be memorised. It shows whether you ask clarifying questions before working, whether you have a repeatable structure, and whether you think past the happy path to misuse, failure and safety — in about five minutes.
How many test cases should I give for a 'how would you test' question?
None, as a number. Give categories, not a count: cover the axes with a line or two each, then ask which area they'd like you to go deep on. Volume of cases is not what is being scored; coverage of thinking is.
What should I ask before answering?
Three or four questions that change the answer: what kind of object exactly, who uses it and where, what it connects to, and what 'correct' means here. Candidates who start listing immediately lose the mark that the question exists to award.
Is it a bad sign if I don't know the product?
No — the interviewer usually picks something you can't have prepared for. That's the point. Not knowing the domain is expected; not having a method for approaching an unknown domain is the failure.
How long should the answer take?
Four to six minutes, structured: about a minute clarifying, three or four walking the axes at one or two lines each, and a close stating your assumptions and priorities. Then stop and invite them to pick a direction.