The GTM engineer case study interview
A GTM engineer case study is a more involved interview format: you hand the candidate a problem they could plausibly face in the job, with a wider scope and more ambiguity than a single question carries. It runs either live, in a 30 to 60 minute session, or as a take-home they spend a few days on, and what you are assessing changes with the format. This page has twelve live exercises, ten take-home briefs, and what needs to go around a brief before either one tells you much.
What is a case study interview for a GTM engineer?
A case study gives the candidate a problem they would plausibly face in the job, and it is deliberately bigger and looser than a single interview question. The scope is wider, some of it is ambiguous, and there is more than one defensible direction to take it in. That is the point. A single question tells you whether somebody knows a thing; a case study tells you how they think when the problem does not come with edges.
You can run it two ways. Live is usually 30 to 60 minutes, with the candidate working the problem while you supply extra context when they ask for it and push with follow-ups. A take-home is a few days on their own.
What you are assessing changes with the format. In a live exercise you are mostly testing whether they can think on their feet, sit with an ambiguous problem, and ask the questions that narrow it down. In a take-home you are looking at the finished work, though these days you also need follow-ups in the session afterwards to work out how much of it they actually did and how much a model did for them.
Neither round settles much on its own. The five rounds covers what each one is good for. This page is the practical part: which of the two to run, what goes in the brief, and how to read what comes back.
Which one should you run?
Neither is better. They show you different things, so the question is which one this role actually needs. Four things to think about.
-
01
Can the work be done in an hour?
Some of it cannot. Standing up sending infrastructure, running three vendors against real credit until the file is enriched, labelling an evaluation set: an hour-long version of any of those measures typing speed and nothing else. Other work is really one decision, and the decision comes early. Which vendor goes first, or what a classifier should do with the records it cannot read, gets settled in twenty minutes and argued about for the rest of the hour. That kind fits a live round.
-
02
Does watching them work tell you anything?
Usually, yes. You get to see how someone opens a problem they have never seen before. Give two candidates the same broken funnel and one starts naming tools while the other starts asking which step the drop is at. That difference is real, and it only exists while it is happening — whatever either of them writes up afterwards will not show it.
The catch is that thinking out loud is its own skill. Someone who goes quiet and comes back with a considered answer may well be the better hire, which is the case for a take-home. The case against is that real work rarely gives you all the time you want, so if this role gets asked for a call before the analysis is finished, an hour of watching that is worth having.
-
03
How much time will candidates actually give you?
A take-home asks for more of the candidate's time than a live exercise does. The candidates you most want are usually employed and have other processes running, so they have the least time to spare and the least reason to spend it on you, which means a heavy brief filters hardest against the people you were trying to hire. Make it big enough to tell you something and small enough that a good candidate does not walk, say up front how long it should take, and use none of what comes back.
-
04
Can you tell a good design from a good presentation?
A take-home almost always ends with the candidate walking you through it, so both rounds get scored by whoever was in the room. The difference is what is in front of you. A live exercise produces something rough that you watched get made. A take-home produces something finished, and finished work is easier to be impressed by than to judge. If nobody on the panel has run these systems, a take-home sorts your finalists by how well they present.
The live exercises
These are exercises you run with the candidate in the room. Each one comes with its own numbers already in it, so they can start on the actual problem instead of spending the first ten minutes working out what they are allowed to assume. All of them are answerable inside an hour.
Ask the candidate to talk through what they are doing as they do it, since that is most of what you are there to see. When a follow-up asks what they are assuming, ask it while they are still working rather than saving it for the end.
Twelve is more than you would use in one round. Filter to the capability domains this role actually owns and pick the two or three that fit the hour you have. Each one opens up to show the follow-ups to ask mid-exercise, and what strong and weak work looks like on that particular exercise.
Filter by what the role owns
-
01 Reply rate on our main sequence sat just over three percent and halved three weeks ago. Here is the sending dashboard for the last quarter and the change log for the four sending domains. The list source changed a month ago, which is the explanation everyone here has settled on. Work out what actually happened. Whether the timeline gets checked before the offered cause
Follow up with
- What date does the drop actually start on?
- Which of the other numbers on that dashboard moved with it, and which held?
- What would you do tomorrow morning, before you are certain?
Strong answer
A strong candidate lines the dates up before doing anything else, and they do not fit: the list changed a month ago and reply rate held for a week afterwards, which is a week longer than a worse list takes to show. From there they work the change log towards that week and find the new sending tool that went onto the domains without the authentication being updated to match. The dashboard settles it, because bounce rate barely moves while placement falls, and a strong candidate reads those two against each other before naming a cause.
Weak answer
A weak candidate takes the list change because it is the only cause anyone has offered, and spends the exercise on data quality. The week between that change and the drop is on the screen in front of them and never gets looked at. What comes back is a proposal to roll the list source back, which costs a fortnight and fixes nothing.
-
02 Here is one company's website and our ideal customer profile, which has six criteria in it. Use whatever AI you normally use and give me a fit score for that company in twenty minutes. Talk me through it while you work. Whether scored and guessed stay separate when evidence runs out
Follow up with
- What has the model just told you that you have decided not to believe?
- Which of the six did you find real evidence for on the site?
- What would you check before letting this run across five thousand accounts?
Strong answer
A strong candidate works criterion by criterion instead of pasting the whole site in and asking for a number. A marketing site carries evidence for two or three of the six at most, and the useful part of the twenty minutes is watching them mark the rest unknown rather than letting the model fill them in. They read the output back against the page, so when it names a customer or a headcount the site never mentions, they catch it and say so out loud. At twenty minutes you get a score plus the criteria it could not reach, still listed and still unscored.
Weak answer
A weak candidate pastes the site into a chat window, asks for a fit score out of ten, and reads the number back with the model's reasoning attached as though it were their own. The six criteria never get taken one at a time, so nothing separates what the page said from what the model supplied.
-
03 Design the enrichment waterfall for this ICP against these three vendors. They charge one, two and five cents a match, we run about thirty thousand records a month, and you have four cents a record to spend. Tell me what you drop. Whether the cost ceiling is arithmetic or a number they nod at
Follow up with
- What are you assuming the first vendor matches on?
- What does the five-cent vendor have to be right about to earn a place at all?
- How would you know a month in that you ordered them wrong?
Strong answer
A strong candidate does the arithmetic out loud. Four cents a record against a five-cent vendor means that vendor cannot be a tier everyone passes through; it only ever sees what survives the first two, and only if those came in far enough under the ceiling to pay for it. So the cheap vendor runs on everything, the middle one on its misses, and the expensive one on a slice they can name and defend. Expect them to answer the last part properly: the records nobody will work this quarter and the fields no downstream system reads both come out of the design, and they will say which.
Weak answer
A weak candidate orders the three by match rate, puts the five-cent vendor first because it is the strongest, and arrives at a per-record cost above the ceiling without noticing. Asked what they would drop, the answer is a shorter list, which shrinks the volume rather than the unit cost the brief set.
-
04 Here are thirty account records, six of which have almost nothing in the description field. Write the prompt that classifies each one against our ICP, run it over all thirty, then show me where it fails. Whether the prompt has an answer for input it cannot classify
Follow up with
- What is your prompt doing with the six thin ones right now?
- How do you know the ones it got right were not luck?
- What would you change if it had to run over fifty thousand of these?
Strong answer
A strong candidate looks at the thirty before writing anything, finds the six thin ones, and builds the prompt so it can come back saying it does not know. That single decision is most of the exercise: a classifier with two outputs will guess on those six, confidently, and the guesses arrive looking exactly like the answers. When the run finishes they read those six first. The failures they show you name the record and the line in the prompt that produced the wrong call, and one of them is usually a case it got right for the wrong reason.
Weak answer
A weak candidate writes the prompt from the ICP alone without reading the records, runs it, and reports the misses as a rate. The six thin records come back classified as confidently as the rest and nobody notices. Asked where it fails, the answer is that the prompt needs work, with no record named and no reason given.
-
05 An LLM scores every inbound lead we get, about nine hundred a month, and roughly forty of those become opportunities. Show me how you would find out whether the scoring is any good. Use whatever tools you normally use. Whether a scorer gets measured or read for plausibility
Follow up with
- Where are the labels you are using coming from?
- Which direction is your own judge wrong in, and how would you find that out?
- What would you run before a prompt change ships?
Strong answer
The forty is what a strong candidate goes after. At forty opportunities in nine hundred leads, a sample drawn at random is almost all negatives and says nothing about the scores that mattered, so they stratify: everything scored high, plus a sample of what was scored low. The labels are usually already in the CRM as what became of the lead, which beats hand-labelling, and they will say why. Watch whether they look at the low scores at all, because a scorer that is wrong about who it rejected is the expensive failure and it never generates a complaint.
Weak answer
A weak candidate pulls a sample out of last month's nine hundred, agrees with most of the scores, and calls the scorer reasonable. Asked what would count as it being wrong, there is no answer, because nothing was compared against what happened to those leads afterwards. The forty opportunities are in the brief and never become the labels.
-
06 Here are forty account rows out of our CRM. Subsidiaries, trading names and domain variants are all in there, and nine of the forty have open pipeline on them. Merge the list down and tell me the rule you used. Whether merging is treated as matching or as a commercial call
Follow up with
- Which two are you least sure about?
- What did you do differently with the nine that have pipeline?
- Where does that rule break on the next forty?
Strong answer
A strong candidate sorts the forty into buckets before merging anything: the same company under a different domain, a trading name, and a subsidiary that may genuinely buy on its own. The last bucket does not get merged, it gets a parent. The nine rows with open pipeline are where to watch them slow down, because a merge that moves an opportunity onto somebody else's account is the mistake that costs a real person money. Expect those nine to go to a human rather than get decided in the room. The rule they state at the end is one per case rather than a single match threshold.
Weak answer
A weak candidate matches on how similar the company names look, merges everything above whatever threshold seems reasonable, and hands back a shorter list. Subsidiaries disappear into their parents with nothing left to show they were ever separate, and the nine rows with pipeline get treated exactly like the other thirty-one.
-
07 Here are forty inbound leads from last week and the routing rules that are live today. Six of them reached the wrong rep. Find those six and change the rules so it stops happening. Whether six failures get one cause or six exceptions
Follow up with
- What do the ones you have found so far have in common?
- What does your change do to the other thirty-four?
- Which of your new rules would you be nervous shipping on Monday?
Strong answer
A strong candidate reads all forty against the rules before changing anything, because the six are not six problems. They share a cause, and once it is named the fix is one rule rather than six conditions. The second half is where people stop early: a rule change applies to everything the router sees, so a strong candidate runs the new rules back over the thirty-four that were right and checks they still are. Listen for them asking what happens to a lead that matches two rules, because the answer decides whether their fix works at all.
Weak answer
A weak candidate finds two or three of the six, patches each one with its own condition, and hands back a rule set longer than the one they started with. The thirty-four that routed correctly never get re-checked, so nobody knows what the patches cost.
-
08 Ask a model to design the lead-routing rules for a company like ours: eleven reps, about twelve hundred inbound leads a month, and roughly a third of those arriving at accounts that already have an owner. Then pull what it gives you apart in front of me. Whether a confident plan in their own domain gets audited
Follow up with
- Which of its assumptions only fails once real data reaches it?
- What did it leave out completely?
- What would you need to know about the eleven reps before you could route to any of them?
Strong answer
A strong candidate goes after the gaps rather than the wording. Models tend to write routing plans for net-new leads, so the third that land on accounts somebody already owns usually get no rule at all, and that is the first thing to listen for. Expect the rest to follow: no fallback when the assigned rep is away, round-robin that ignores who is already carrying more than they can work, nothing that says how you would find out the routing was wrong. The tell is that they prompt the model and then stop trusting it, reading the plan against eleven named people rather than against routing in general.
Weak answer
A weak candidate reviews it the way you review a document, approving the structure and tightening the language. The plan never says what happens when nobody is available and never mentions the accounts that already have owners, and neither absence gets raised. Some will prompt the model to critique itself and read that back to you.
-
09 The nightly sync into our accounts table ran twice last night. Here is the table this morning: about ninety thousand rows, where yesterday there were seventy-nine thousand. Get it back to one row an account without dropping anything real. Whether they fix the rows or the write that produced them
Follow up with
- What are you keying on, and where did that key come from?
- How do you know one of those extra rows is a duplicate rather than a second real record?
- What stops tomorrow night doing this again?
Strong answer
A strong candidate settles what makes two rows the same before deleting either, and it has to be the source system's identifier rather than every field matching, because a record edited between the two passes differs from its own copy. Eleven thousand extra rows against seventy-nine thousand also says the second pass was not a straight repeat of the first, and it is worth seeing whether they notice that and go looking for the difference. Keeping the later row per key is usually right and they will say why. Expect a senior candidate to go past the clean-up: the same job run twice should have landed the same rows, so the real fix is an upsert on that key and this morning's table is the symptom.
Weak answer
A weak candidate writes a delete against rows whose fields all match, which leaves every pair where something changed between the passes and removes rows that were never duplicated. The job itself goes untouched, so a second double run tomorrow produces the same morning again.
-
10 Here is our inbound funnel from demo request through to opportunity. It runs across a form, the CRM and a calendar tool, and the only thing anybody can measure today is how many opportunities came out: about nine hundred requests go in a month and around seventy opportunities come out. Define the events you would add across those three systems, then name the one number you would put in front of the CRO every week. Whether the identifier gets solved before the events get listed
Follow up with
- How does a booking in the calendar tool get tied to the form fill it came from?
- Which of your events could you not add after the fact?
- What does the CRO do differently when your one number moves?
Strong answer
A strong candidate goes at the join before the events. Three systems means three identities for the same person, and events that cannot be tied together measure three funnels rather than one, so the first thing they want is an identifier that survives the form, the CRM record and the booking. Once that exists, the events they list have to account for the eight hundred and thirty who never became an opportunity, which is the only way nine hundred in and seventy out locates anything. The weekly number they land on is one the CRO can act on, and they will say what decision it changes.
Weak answer
A weak candidate lists a dozen events, one per screen, and never says how any of them attach to the same person across the three systems. The one number comes back as demo requests, which needs no instrumentation to count and moves for reasons that have nothing to do with the funnel.
-
11 Build the trigger that sends an onboarding drop-off email. These are the nine fields our product writes into the CRM, and none of them records when an account was last active. About two thousand accounts sit in onboarding at any one time. Work with the nine. Whether a missing field stops them or gets worked around
Follow up with
- Which of the nine are you using, and what are you assuming it means?
- What does your trigger do to the two thousand accounts that are already mid-onboarding when it goes live?
- Who gets this email who should not?
Strong answer
A strong candidate reads the nine fields first and works out which of them moves over time, because drop-off has to be built out of something carrying a date. Absence of movement in a step field is the usual construction, and a good answer says how long an account has to sit still before it counts, with a reason for the number rather than a round one. Expect them to raise the backfill without being asked: two thousand accounts are already mid-onboarding, and a trigger that evaluates history on the day it goes live mails all of them at once.
Weak answer
A weak candidate asks for a last-active field, is told there is not one, and designs around it anyway. Or the trigger fires on a status the product sets once and never updates, so it either never fires or fires for everybody. Either way the two thousand accounts already in flight go unmentioned.
-
12 Two people in our ops team spend about three hours every Monday working out why the open pipeline number in the CRM and the one on finance's spreadsheet disagree. The gap has sat around fifteen percent since the stage mapping changed in March, and the Monday export drops any opportunity edited after the snapshot ran. Build the smallest thing that stops them doing that by hand. You have an hour and whatever tools you use day to day. Whether they settle which number is right before automating it
Follow up with
- Which of the two numbers is your tool going to publish?
- A rep edits an opportunity an hour after the snapshot runs. What does your tool show on Monday?
- Who changes it when finance redefines what counts as open?
Strong answer
A strong candidate does not automate the Monday compare, because a tool that reproduces a fifteen percent gap every week has saved nobody anything. The two causes in the brief behave differently and they separate them: a stage mapping change moved the CRM's number for every prior week as well as this one, while a snapshot that drops late edits loses a different set of rows each Monday and never the same ones twice. Only the second leaves the two numbers genuinely out of step. What gets built is then small: open pipeline read out of the CRM on the definition finance actually uses, counting the rows the export was dropping, posted where the ops team already looks. Watch what they do with the parts the hour will not cover, because a strong candidate names them out loud and leaves the thing working without them.
Weak answer
A weak candidate automates the Monday compare: two queries, a difference column, and a post showing the same fifteen percent gap those two people were already working out by hand. Both causes are stated in the brief and neither gets used, so the tool never says which number is right and somebody still has to. Asked what it publishes, the answer is both.
No question carries that tag.
The anatomy of a brief
Each take-home below is only the challenge statement. Everything else in a brief is the same document every time you send one, so it is set out here once rather than repeated ten times over. Five parts go around the challenge. The first four are the obvious ones; the fifth is the one almost nobody writes down.
- Scope. What is in the exercise and what is stipulated: which systems they may assume exist, what credentials and sample files come with it, and which parts of the problem you have decided not to ask about. Leave the boundary open and submissions come back at wildly different sizes, so you end up comparing how much time each candidate had.
- Deliverables, named as objects. A one-page memo, a script that runs against the attached file, five slides, a sheet with a column per criterion. Name the format and the length, because a candidate asked for a recommendation guesses at both, and the guesses vary more than the work does.
- The time bound as a norm rather than a deadline. "We generally look to receive this within four days, and tell us if you need longer" costs you nothing to write and returns what the candidate can do instead of what their week allowed. A hard date mostly measures who had a quiet Tuesday.
- What you are evaluating. Three or four dimensions, settled before the brief goes out and printed in it. Sharing them does not let a weak candidate fake them, and it stops a strong one spending Saturday on the dimension you care least about.
- What you are not evaluating. That visual polish is not scored, that the code does not have to be production-grade, that you will read one page and no more, that you do not want the problem statement questioned. This is the part that almost never appears, and it changes what comes back more than any of the other four.
The take-home briefs
Ten challenge statements. Each one is the paragraph that changes from brief to brief, and everything in the five parts above stays the same whichever you send.
They are written at the size this work actually runs at, which is bigger than a screening exercise should be. Cut one down before you send it: drop a deliverable, shrink the file, tell them which half not to build. Whatever survives the cut needs to still be the part you cannot learn any other way.
The follow-ups on each belong to the session after the submission lands, not in the brief itself. Send the challenge statement and keep the follow-ups for the conversation.
Filter by what the role owns
-
13 You have a segment of about four thousand accounts, a product we will brief you on, and no sending infrastructure at all: no domains, no inboxes, no sending tool, no reputation anywhere. Our own domain carries everyone's mail and is not available to you. Design the program, stand up what it needs, and hand over what you built. Nothing in it may assume a domain bought this week can carry volume next week. Whether the warmup sets the schedule or is noticed at the end
Follow up with
- How many inboxes did you land on, and what number did you get there from?
- What is in week one that is not in week six?
- Who answers the replies once this is running?
Strong answer
The ramp is the schedule, and a strong candidate builds backwards from it. Four thousand accounts across a sequence of a known length gives a daily send, a daily send no inbox can safely exceed gives the inbox count, and the inbox count gives the domains, which is where the money goes. Because a new domain has to warm before it sends, the first real send lands weeks out, so they buy domains in the first week and fit the list, the copy and the sequence structure inside that window rather than after it. Authentication is set per sending domain before anything leaves, the sending domains are lookalikes that redirect somewhere real so a prospect who clicks does not hit a dead host, and the ramp is written as a per-inbox daily number by week. A senior candidate also names who handles replies and what they do with an out-of-office or a referral, because a program at this volume produces both from the first week it sends.
Weak answer
A weak candidate picks a sending tool first and treats everything else as configuration inside it. The number of inboxes comes from a monthly send target with nothing behind it, so nobody can say whether four domains or twelve is right. Warmup appears as a task running alongside a launch date that does not account for it, and the handover is a tool account with a sequence in it and no record of what was configured where.
-
14 The file holds two thousand records for our ICP, each with a company domain and a person's name and title. Build a working waterfall across the three vendor APIs whose keys are attached, and report the coverage you reached and how accurate that coverage is. The credit sitting on those three accounts is the entire budget, and it does not stretch to running every vendor over every record. Whether accuracy was measured or taken from the vendors
Follow up with
- Which vendor would you drop if you had to keep two?
- An address came back for a record. What made you believe it?
- What did the records it could not match have in common?
Strong answer
A strong candidate treats coverage and accuracy as two measurements taken two different ways. A vendor returning an address is coverage, and every vendor counts its own generously; whether that address reaches that person is a question no vendor answers about itself. So a verification step goes in, and a domain that accepts anything sent to it comes back marked unverified rather than counted as a match. The budget bites on the checking rather than on the matching, which is where people get it backwards: spend the whole credit finding addresses and there is nothing left to find out whether any of them are real. A strong candidate holds some back, or verifies a sample instead of the whole two thousand, and says which of those they did. Expect the domains to be normalised before a vendor is called at all, because a fair share of two thousand of them arrive as a marketing site, a redirect or a parent company, and a vendor asked about the wrong domain still answers.
Weak answer
A weak candidate runs all three vendors over all two thousand records, exhausts the credit, and reports the union of what came back as the coverage and the accuracy at once. Nothing in the submission separates an address that was checked from one that was merely returned. Asked which vendor earned its place, the answer is whichever returned the most rows, and nobody checked those rows.
-
15 Build an agent that takes an inbound form fill, which is a name, an email address, a company and one free-text line about what they want, and decides whether it fits the ICP in this brief, with a reason a rep can read and check. Two hundred real form fills are in the file. Build the eval set that says how well it does, and write down where it is wrong. It has to return a decision for every record, including the ones where the form was filled in badly. Whether the eval set exists before the agent is tuned
Follow up with
- Which records did you hold back, and when did you label them?
- Show me a reason it produced that you would not put in front of a rep.
- What would have to be true in production before it decided more on its own?
Strong answer
A strong candidate labels a held-back slice of the two hundred before tuning anything, because an agent tuned against every record in the file has no honest number left to report about itself. They read the two hundred first, and those two hundred are messy in ways that decide the design: personal email addresses, one company typed four ways, and a free-text line that is blank about a third of the time. Each of those needs a decided behaviour rather than whatever the model does when nobody specified one. The reasons that come back are grounded in something the agent looked up and can be checked against it, so a rep reading one can tell whether anything was found. Where it is wrong is broken out by the classes of record it struggles with, since a single accuracy figure over the two hundred hides the badly filled ones inside it. Expect the submission to name what would be watched in its first weeks live.
Weak answer
A weak candidate writes a prompt, runs the two hundred through it, and computes accuracy against labels written after reading the output. The reasons restate the ICP in the company's own words, so nothing in them says what the agent actually found. Records with a personal email address or an empty free-text line come back decided as firmly as the rest, and where it is wrong is answered with a paragraph about the limits of language models.
-
16 Here is the CRM schema our lead router keys on, and six months of complaints from reps about leads reaching the wrong person, a hundred and forty of them with the lead record attached. Rewrite the rule set, and show what your version does with each of the hundred and forty. The rules have to run on the fields in that schema as they stand: you do not get a new field and you do not get a clean one. Whether the complaints are read as a sample or as everything
Follow up with
- How many misroutes do you think happened that nobody wrote up?
- Which of the hundred and forty does your version still get wrong?
- What would you watch for a month after this ships?
Strong answer
A hundred and forty complaints are the misroutes somebody was annoyed enough to write up, and a strong candidate says so before rewriting anything, because a rep who gave up filing them produces no evidence at all. Reading the schema against the records then splits the pile: a good number of these were never rule failures. The field the rules key on is free text a rep types, or it arrives empty on anything the form created, and no rule written over an empty field could have routed those leads anywhere. That half is a data problem and gets named as one. What comes back is short, because six months of complaints usually collapse into three or four causes, and it is walked through the hundred and forty case by case, including the ones the new rules would still send to the wrong person and the ones the old rules got right by luck.
Weak answer
A weak candidate treats the hundred and forty as the specification and writes a condition for each pattern in them, so the rule set grows by everything anyone ever noticed. Leads nobody complained about are never run through the new rules, so what the rewrite costs is unknown on the day it ships. Complaints caused by an empty field come back with a rule written over that same empty field.
-
17 You have the object and field layout of the CRM we are on, with row counts and fill rates per field, and the layout of the one we are moving to. Produce the field map, the plan for moving, and the way back if it goes wrong. Sales does not stop selling while this happens, so there is no window in which both systems are idle. Whether the plan survives records changing while it runs
Follow up with
- Which object do you cut over first, and which last?
- A rep closes a deal at 9am, mid-window between that object's 2am export and its afternoon cutover. Where does the deal live an hour later?
- What has to be true before you turn the old system off?
Strong answer
The absence of a quiet window is the whole exercise, and a lead-level plan is shaped by it: history loaded in bulk, then a period where both systems are live and changes flow one way only, with a cutover moment named per object rather than one for the migration as a whole. Fill rates get used rather than read. A field populated on four percent of rows does not need a home in the new schema, and saying which fields die is most of what keeps the target clean; fields with no counterpart come back as decisions for a named person instead of being mapped to the nearest thing available. Open opportunities and their owners move first, because that is what a rep notices missing within the hour. The way back is where people separate. A rollback that means the old system stays authoritative until stated conditions hold is a rollback; restoring a backup is not one, because the old system has gone on taking writes the entire time.
Weak answer
A weak candidate produces a field-by-field map and a cutover weekend. Every source field gets a destination, including the ones nobody has filled in for two years, so the new system opens with the old one's mess already in it. The plan assumes the source stops moving when the export runs, and the way back is the backup file. Nothing in it says what happens to a deal edited during the load, or which object goes first.
-
18 Pick three public signals that say a company in the supplied ICP is worth reaching this quarter, build the pipeline that detects them across the eight hundred accounts in the file, and report the precision of each one against the labelled sample sitting in there with them. Two things are fixed: a signal firing on more than a fifth of those accounts in a month is noise, and the pipeline has to run again next month without you in the room. Whether precision is computed against the labels or asserted
Follow up with
- Which of the three would you switch off, and what did finding that out cost?
- What happens the second month to a company that fired in the first?
- Where does your precision number come from, record by record?
Strong answer
Which three they pick tells you most of it. Hiring, funding and headcount are usually where people start, and the version of each that survives contact with a labelled sample is narrower than the one people reach for first: a company hiring the role that owns the problem the product solves, rather than a company hiring. A strong candidate states what each signal is supposed to imply and then checks whether it does, which is what the labels are for, and precision comes back per signal with the numerator and the denominator shown. Expect one of the three to fail that check and be reported anyway, usually the one firing on companies that are already customers or already in pipeline. The monthly part gets a real answer too: one funding round reaches three sources under three dates and has to collapse to one event, and something has to remember what already fired so March does not arrive again in April.
Weak answer
A weak candidate picks three things that sound like intent, matches keywords over whatever each source returns, and reports how many accounts fired. Precision is claimed rather than calculated: the labelled sample is in the same file and never gets joined to the output. The pipeline runs once, by hand, over a list somebody already filtered, and nothing in it would stop one event being counted three times or counted again next month.
-
19 Demo-to-opportunity conversion sits at nineteen percent and the board wants it above thirty. The file has fourteen months of demos with the account on each one, the segment and the sequence it came from, and what happened to it afterwards. Find where the conversion is being lost and fix it. Whatever you propose has to be runnable next quarter with the team as it stands. Whether conversion gets cut by segment before it gets fixed
Follow up with
- What do the accounts that converted have in common that the rest do not?
- What does your fix do to the number of demos booked next quarter?
- Which segment would you stop selling to, and who has to agree to that?
Strong answer
A strong candidate takes the nineteen percent apart before believing in it. Grouped by segment it stops being one number: the accounts that resemble the customers who stayed convert at a rate nobody would have escalated, and the volume sits in the segments that convert close to nothing. Averaged together those produce nineteen, which means every hour spent on the demo itself is spent at the wrong stage. The work then runs backwards from the accounts that did convert, describing them in terms the file can express: size, what they already run, who took the meeting, which sequence and which segment they arrived through. That description is the ICP, and nobody wrote it down when the sequences were built. What comes back changes who gets a demo rather than what happens inside one, and the honest version says the demo count will fall as the top of the funnel narrows, and says it before anyone sees the drop.
Weak answer
A weak candidate accepts the demo as the place conversion is lost, because that is the stage the brief names, and works on the demo: a tighter script, a discovery template, a follow-up sequence, a scorecard for the reps. Segment sits in the file as a column and the analysis never groups on it, so nineteen percent stays one number from start to finish. The proposal would raise conversion only if every segment behaved like the average, and none of them does.
-
20 This file has eighteen months of touchpoints joined to closed-won and closed-lost opportunities. Three things are true of it: paid clicks carry a click id and nothing else does, everything before a visitor accepted the cookie banner is missing, and the first four months predate the current UTM convention, so one campaign appears under three names. Build the model, and write the caveats that have to travel with its numbers when a CRO reads them. All of it has to be reproducible from that file: no cleaning by hand that cannot be run again. Whether the gaps end up in the model or in the caveats
Follow up with
- Which channel does your model flatter, and what gave it away?
- What decision would you stop a CRO making on these numbers?
- What would you instrument now so the next eighteen months are cleaner?
Strong answer
Most of the work goes into what is missing rather than into the weighting, because any scheme applied to a partial record is confident about the wrong channel. The three gaps are not the same kind of gap, and a strong candidate sorts them before choosing any weighting at all. The campaign naming can be mapped and the mapping re-run, so it costs an afternoon. The pre-consent touchpoints cannot be recovered at all, and their absence quietly moves credit to whatever is still visible, which is usually direct and organic. The click id is the finding worth having: paid can prove a touch happened and no other channel can, so any model that rewards a matched touch pays paid first, and that is a property of the instrumentation rather than of the channel. What reaches the CRO says what the model is good enough to decide, which is where to move budget at the margin, and what it cannot settle at all. The caveats sit beside the numbers rather than in an appendix.
Weak answer
A weak candidate reaches for a shape, first touch or last touch or an even split, applies it to the rows as they arrive, and reports revenue by channel to the dollar. Direct comes out large and gets read as brand strength rather than as the shadow of the events nobody recorded. One campaign under three names is counted as three campaigns. The caveats say that data is never perfect.
-
21 Here is the trial-to-paid program as it runs today: nine emails over four weeks, the entry and exit condition on every step, and last quarter's sends, opens and conversions for each. Accounts pile up at step four and most of them never reach step five. Rebuild it, and define what you would measure to know the rebuild is better than what it replaced. The same accounts are being emailed by two other programs and by whichever rep owns them, and none of that stops for your version. Whether the program is rebuilt for the inbox it lands in
Follow up with
- Is step four an email that fails or a condition that never fires, and how did you tell?
- What does your program do when another one is already mid-send to that account?
- What number would tell you in three weeks that this is worse?
Strong answer
The per-step numbers get read before a single email is touched. Opens holding up at step four while nobody advances points at the exit condition rather than at the copy, and opens collapsing there points the other way; last quarter's numbers carry that reading and the emails themselves do not. The other senders are the part most people skip. An account sitting in three programs plus a rep's own sequence receives more mail than any one of them was designed to send, so the rebuild has to state what it does when another sender is already in there: hold, skip the step, or drop the account out entirely. What they choose to measure decides whether the rebuild can be judged at all. Sends and opens per step are what the old program already reported and they compare it only against itself; conversion measured on the accounts that entered in a given week is what says which version is better.
Weak answer
A weak candidate rewrites the nine emails, adds two, and reports that the new copy reads better. Step four is treated as a copy failure without anyone checking whether an account was ever able to exit it, so the pile-up survives the rebuild. The two other programs and the rep go unmentioned, which means the new version lands on top of them, and the measure proposed is open rate, which the old program was already reporting.
-
22 Attached is the revenue team's whole tool bill: what each tool costs a year, whether it is billed by the seat or by usage, the last login date on every seat, and which other tools read data out of it or sign in through it. The total is a little over four hundred thousand, and the data platform at a hundred and thirty of that is where the CFO has asked us to start. Take forty percent out of the bill. Nothing the team named as load-bearing can stop working, and that list is in the file too. Whether the whole bill gets read or the biggest number gets hit
Follow up with
- Where did the forty percent come from, line by line?
- Which of your cuts needs a migration before it can happen?
- What did you leave alone that the CFO expected you to cut?
Strong answer
A strong candidate works the whole list before going near the tool they were pointed at, and the login dates are why. Several of the small tools have most of their seats untouched for a quarter, and together they come to more than any discount anyone would win on the data platform, which the file shows a dozen other tools reading out of, so cutting it is a project rather than a saving. Forty percent will not come out of unused seats alone, so the answer arrives in layers: cancel what nobody opens, size seats to logins, move what is not load-bearing to a cheaper tier, and cost out the one or two consolidations that need a migration before proposing them. A strong candidate also separates the seat-priced tools, where a cut lands next month and can be undone, from the usage-priced ones, where the spend follows a volume somebody else controls and cutting it means changing how the team works. They say which cuts they would not make, because the file names what the team calls load-bearing and a plan that quietly overrides that list is one nobody will carry out.
Weak answer
A weak candidate opens on the data platform because it is the largest number on the page, and spends the work on a negotiation and a tier downgrade that gets nowhere near forty percent. Or the number is reached by cutting the two biggest invoices, one of which three other tools sign in through, which the file states on that tool's own row. Neither the login dates nor the columns saying what reads from what end up compared against anything, so the plan can say how much comes off the bill and not what stops working.
No question carries that tag.
A brief has to say what it is not asking for
Here is a situation that comes up a lot. A company sends a candidate a take-home: here is a leak in our funnel, tell us how you would fix it. The candidate digs into the data and works out that the funnel is not really the problem. The wrong accounts are coming in at the top, and fixing that would mean changing who the company sells to, which is a decision the brief never invited them to touch.
So the candidate has to pick one. Fix the step the company asked about, or hand back a finding about a decision nobody asked them to question. They fix the step, because when a brief does not say otherwise, going outside it feels like inviting a rejection that says "you didn't read the brief."
The company rejected them anyway, and the reason they gave was that the candidate had not been proactive enough.
The candidate had to guess what the company wanted, and that is the company's fault rather than the candidate's. Some hiring teams want somebody who diagnoses past the brief. Others want somebody who solves exactly what was asked. Both are legitimate, most teams feel strongly about which one they want, and almost none of them write it down. So every candidate decides it privately, and the submissions you end up comparing differ on a question you never asked.
There is a second preference that works the same way, and it showed up in the same rejection: how fast you want the work done. The company told this candidate they were not fast enough, while praising the quality of what they produced. If nobody tells a candidate that turnaround is being scored, they will spend an extra day making the work better, and that extra day is what loses them the job.
Both of those are the brief failing to say what it is not asking for, and both are fixed by one sentence. Either "the problem statement is our reading of it, so tell us if you think we have it wrong — we would rather hear that than have you solve around it", or the opposite: "we have already settled what the problem is, so solve the one in the brief." Either sentence is defensible. Having neither is not, because it turns the exercise into a guess about what you want.
So decide what you are not evaluating before you write the brief, and put it in there. That is the harder half to agree on internally, but it is an argument the hiring team is going to have in the debrief regardless. The only difference is whether it happens before a candidate spends their weekend on the work, or after.
Frequently asked questions
What is a case study interview for a GTM engineer?
A larger, more open-ended problem than a single interview question, usually something the candidate would plausibly hit in the job. It runs either live, in a 30 to 60 minute session, or as a take-home over a few days. Live, you are watching how they open the problem and what they ask. With a take-home you are looking at the finished work, and at how they answer for it in the session afterwards.
Is a case study the same as a technical assessment?
A technical assessment is any evidence of building that you collect before making an offer. For this role the case study is the version that works: a bounded slice of the systems the hire will actually own. A coding puzzle rejects builders who do not write much code, and a tool quiz passes people who have only read the documentation. So a GTM engineer's technical assessment is usually a case study, run live or as a take-home.
Should a GTM engineer case study be live or a take-home?
Live when the decision you care about happens in the first few minutes and the work fits in an hour. Take-home when the artifact is the thing the job actually produces, you have somebody who can judge it, and you have booked a session for the candidate to walk you through it. Employed candidates drop out of long take-homes, so keep it to the smallest version that still tells you something.
How long should a GTM engineer take-home be?
A few hours of work. Keep it there by cutting deliverables, not by asking for less care. State the turnaround as a norm: we generally look to receive this within four days, and tell us if you need longer. Send every candidate the same brief so you can compare the submissions, and expect the strongest applicants, who are usually employed, to weigh the cost before they start.
Hiring a GTM engineer?
Are you a GTM engineer?
Get on the radar