AI policy for GTM engineer interviews

Set the AI policy according to the evidence each round is meant to produce. Let candidates use the tools available in the job for case studies and other work samples. In a behavioral round, ask candidates to answer from their own experience without live model assistance. State the policy in the invitation so every candidate works under the same conditions.

Hold with AI open 23
Need proctoring 8
Bank version 2026.7

Should candidates use AI in a GTM engineer interview?

Allow AI when the round evaluates work produced with AI. Require disclosure of the tools used and ask the candidate to defend the result. Close live model access when the round evaluates the candidate's unaided account of prior experience.

The taxonomy includes AI-assisted and AI-engineering approaches. For roles using either approach, a case study should show how the candidate prompts, checks, revises, and integrates model output.

Behavioral questions ask what happened on a system the candidate ran. Preparation is expected, but the live answer should remain the candidate's account. Follow-ups about sequence and specific decisions help test that account.

What each round does about AI

Write the permitted tools, disclosure requirements, and evaluation criteria into each invitation. The policy changes by round because the evidence changes by round.

  1. 01

    Behavioral: no

    The questions ask what the candidate built, what broke, and how they found out. Ask candidates to close live model access, then use follow-ups to examine their account of the work.

  2. 02

    Technical: allowed, with the recall questions anchored

    Allow it and anchor recall questions to the candidate's work. Two questions in our behavioral set ask for a fact any model has: how to set up a cold-email domain that also carries the team's day-to-day mail, and how to re-run a bad enrichment day without duplicating rows or paying twice. Both have a right answer, and it is in the training data.

    Two others ask for facts about the candidate's own systems: the DNS records on the last sending domain the candidate set up themselves, and the per-inbox daily volume the last program they ran carried. The second pair is anchored to a system the candidate operated.

    So keep the recall questions and anchor them. Ask what they set, on which domain, and how they arrived at the number, rather than what the number should be.

  3. 03

    Case study: yes

    The candidate is producing an artifact under the conditions they would use at work, including permitted models. Score what they took from the model, what they threw out, and whether they questioned the brief you handed them or executed it as written.

  4. 04

    Portfolio review: yes

    A portfolio walkthrough covers work that already exists, so there is nothing to police. If a model wrote part of it, that is a fact about the work and worth asking about. Ask which parts it wrote, what the candidate checked before shipping, and what happened to the thing after it went live.

  5. 05

    Reverse questions: does not arise

    The candidate is asking you here, so there is no policy to set. What they ask about AI is still information. Someone who wants to know which parts of the stack you expect a model to run, and who reviews its output, is asking a question you only think to ask after owning one.

Why it depends on how they build

The taxonomy includes four build approaches. Match the exercise and AI policy to the approach named in the mandate.

  1. 01

    No-code automation

    The work here is wiring platforms together (Zapier, Make, n8n, Clay), with no codebase underneath. A model can draft the logic, so an AI-allowed exercise tests whether the candidate knows what that logic does at volume, on a bad record, and when a vendor changes its schema. Ask them to explain a step they did not write.

  2. 02

    AI-assisted

    Here the work gets done through models inside software the candidate already runs — a generated column in a tool they own, or a research pass they never wrote by hand. Taking AI out of a round for a role like this removes the thing you are hiring for. Watch what they prompt, what they keep, and what they hand over without reading.

  3. 03

    AI engineering

    The candidate builds the model into a system that runs without them, with retrieval, prompts under version control, and evaluations that catch a regression before a prospect sees it. Ask how they would detect degraded output and what they would do after an evaluation failed.

  4. 04

    Custom code

    The role expects real programming, in Python or TypeScript or Go, for services and integrations nobody sells off the shelf. A model writes a plausible first draft of most of it, so the exercise is whether the candidate can read the draft. Ask what they changed before it ran, and what they deleted.

The questions that stop working when a model is open

A remote interview cannot fully verify that a candidate has closed every model. These scenario questions become less diagnostic when a model can propose the initial answer because they do not begin with the candidate's own experience.

Use them in a live conversation and follow them with questions about a system the candidate operated. Each entry includes follow-ups that move from a general answer to specific decisions and evidence.

  • 02 A funnel you own is leaking between acquisition and activation. What are the first three things you do? Whether the first reflex is diagnosis or tooling

    Follow up with

    Strong answer

    A strong candidate treats the stated leak as somebody's hypothesis and starts by asking what it is based on. The first move is questions rather than actions: which segment is leaking, whether it has always leaked or started recently, and whether the activation event is instrumented well enough to say. They usually get to whether the people arriving were the right people at all, which moves the problem out of the funnel and into targeting.

    Weak answer

    A weak candidate accepts that the leak sits between acquisition and activation because the question said so, and answers with tools they would add: a session recorder, a warehouse, a dashboard. Nothing in the three moves would tell them the premise was wrong.

  • 03 You are handed a cold-email domain that also needs to run the team's day-to-day mail. How do you set it up? Whether deliverability knowledge is operational or repeated

    Follow up with

    Strong answer

    A strong candidate stops on the premise, because a domain's MX records point at one provider and the setup as described does not exist. They separate the sending domain from the corporate one, usually a lookalike that redirects, and only then talk about authentication and the ramp. Ask for the ramp week by week and they give you weeks and numbers rather than a principle.

    Weak answer

    A weak candidate takes the premise at face value and starts configuring inboxes and warmup on the corporate domain. The conflict is sitting in the question, and reciting a sound warmup ramp on top of it does not make the domain safe to send from.

  • 23 Your CRM says a campaign produced 40 opportunities. The attribution report says 12. Where do you look? Whether they reach for definitions before dashboards

    Follow up with

    Strong answer

    A strong candidate goes at the definitions before either system, because both numbers can be right under their own rules: which object counts as an opportunity, what window, first touch or multi-touch, how duplicates were merged. They will ask what the number is for before choosing which one to take to the board. A good answer ends with the rules written down where both teams read them, since this argument returns every quarter.

    Weak answer

    A weak candidate assumes one of the two systems is broken and starts checking the integration. A gap between 40 and 12 is wide enough to be a definition rather than a sync failure, and that possibility never gets raised.

  • 24 A vendor returned garbage for yesterday's enrichment run and you need to re-run it. How do you do that without duplicating rows or paying for the same records twice? Whether re-running a job is routine for them or something to avoid

    Follow up with

    Strong answer

    A strong candidate replaces the affected day rather than appending to it, and makes the write idempotent on a stable key so running it twice lands the same rows. Ask what happens when the re-run dies halfway and the answer is that they start it again, which is the point of having built it that way. Expect a concrete check that the day is now right, usually a count plus a sample compared against what the vendor returned.

    Weak answer

    A weak candidate deletes rows by hand, or re-runs the job and plans to clean up the duplicates afterwards. The question names both costs it is protecting against, and paying the vendor a second time for records already bought does not come up.

  • 25 A sequence that was replying at four percent drops below one in a week. Nothing about the copy changed. Walk me through it. Whether they reach for infrastructure or for rewriting

    Follow up with

    Strong answer

    A strong candidate uses the sentence about the copy, and treats copy as the last suspect precisely because it did not change. They go at delivery first: authentication, domain reputation, a seed test, and whatever else moved that week in volume, provider or list source. Expect them to hold the answer open, since an unchanged-copy collapse can be placement, a worse list, or a shift in who is being mailed, and a week is not long enough to have separated those.

    Weak answer

    A weak candidate starts rewriting subject lines, which the question rules out in advance by saying nothing about the copy changed. Or they blame list quality without first establishing whether the mail is arriving at all.

  • 26 Inbound demo requests are reaching a rep 30 hours after they come in. The SLA is five minutes. Where do you look first? Whether they can decompose a routing path in order

    Follow up with

    Strong answer

    A strong candidate walks the path in order and says it out loud: form submission, record creation, assignment rule, queue, notification, whether a rep was awake. Thirty hours against a five minute target is a stop somewhere rather than slowness everywhere, and they are hunting for the step where the clock parks. What they want afterwards is a timestamp at each hop, so the next slip shows up as a number rather than as a complaint.

    Weak answer

    A weak candidate proposes a new routing tool, or decides the reps are ignoring their alerts, before knowing whether the lead reached one. The gap in the question is wide enough to locate precisely, and nothing in the answer would locate it.

  • 27 You are four days into a test of a new outbound sequence, the variant is replying 30 percent better, and a VP wants it rolled out to the whole list on Monday. What do you say? Whether they can hold a line on evidence without stonewalling

    Follow up with

    Strong answer

    A strong candidate explains in plain language why a four-day read moves around, without leaning on the word significance, and goes to what was agreed before the test started, including the case where the answer is nothing. Then they give the VP a route: what they would need to see to call it on Monday, or a partial rollout that keeps the test alive. Expect them to say what would make them happy to call it early, because sometimes the gap is wide enough and the downside small enough.

    Weak answer

    A weak candidate concedes on the spot because a VP asked, or lectures about sample size and offers nothing the VP can act on. Either way the person who has to decide on Monday leaves with the options they walked in with.

  • 28 A prospect receives four emails in one day from four different systems. How do you stop that happening again? Whether they fix the instance or the architecture

    Follow up with

    Strong answer

    A strong candidate reads four systems as the problem rather than four emails, and puts the decision to send behind one arbiter holding suppression and frequency for all of them. They will name which system that should be and why, usually the one every tool already syncs through. The harder half is the second follow-up, and a good answer says how the next tool somebody buys inherits the rule instead of routing around it.

    Weak answer

    A weak candidate turns off the offending campaign, or proposes a standing meeting between the teams that own the four systems. The question asks how it stops happening again, and a fifth system could be bought next month under either answer.

Design take-homes for AI access

Assume candidates can use models to interpret the brief, draft the artifact, and critique it before submission. State whether that use is permitted and require a short account of the tools and process used.

Add a live review where the candidate explains the architecture, identifies model output they changed or rejected, and responds to a new constraint. Score the submitted artifact and the defense against criteria shared before the exercise.

Frequently asked questions

Should candidates be allowed to use AI in a GTM engineer interview?

Allow it when the round evaluates work the candidate would produce with AI on the job, and ask for disclosure of how it was used. For behavioral questions about prior experience, ask the candidate to close live model access and answer from their own account.

How do I stop candidates using AI in a remote interview?

You cannot verify it completely. State the rule, apply the same conditions to every candidate, and design follow-ups around the candidate's own work: what they configured, on which system, at what volume, and what happened after launch.

Does an AI-assisted take-home still tell me anything?

Yes. Score the artifact, the candidate's process, and their live defense. Ask what they questioned in the brief, what they changed or rejected from the model, and how the design responds to a new constraint.

Which interview questions still work if the candidate has ChatGPT open?

The ones anchored to what this candidate personally did. Of the 31 behavioral questions in our bank, 23 remain anchored to the candidate's own work, because they ask for specifics from the system they ran: the DNS records they set, the per-inbox volume they ran at, and what broke after they left. The other 8 put a scenario to the candidate and become less diagnostic with model access.

Hiring a GTM engineer?

Are you a GTM engineer?

Get on the radar