The two ways to use ChatGPT in support (copilot vs. frontline)
There are two ways to put AI in support and they're radically different. Mixing them up — or pitching one as if it were the other — is the main source of complaints and tanking CSAT.
| Mode | What it does | Risk |
|---|---|---|
| Copilot for the human agent | Suggests responses; the human decides | Low — the human filters |
| Frontline (visible to customer) | Replies to the customer directly | High — errors are visible |
When frontline IS the right call (strict criteria)
Frontline only if you meet all 4 conditions — not 3 of 4:
- Solid, current knowledge base structured for RAG.
- Working human escalation with context that carries over.
- Volume that justifies the investment (>500 repetitive tickets/month).
- You accept CSAT drops 5-10 points the first month while you tune.
Prompts and templates for the copilot
For the agent copilot, the highest-ROI prompts:
- Customer history summary. "I'm sending you the history. Return: 1) what happened before, 2) what's still unresolved, 3) what tone to use."
- Response draft. "Customer is asking [X]. Applicable policy: [Y]. Generate a response: empathetic, clear, 3 paragraphs max."
- Tone suggestion. "Frustrated customer [context]. Suggest 3 alternative openings that acknowledge the frustration before explaining the solution."
- Technical translation. "This is what the customer said. This is what they probably mean. Here's the reply in customer-friendly language."
When to jump to a dedicated 24/7 AI Support system
- Your support team takes longer to review the suggestion than to write it themselves.
- More than 30% of tickets are repetitive and still escalate to the team.
- Volume grows faster than the team and you can't hire fast enough.
- You need 24/7 coverage and it's not viable with humans.
In any of those four cases, move to a dedicated AI Support system — not raw ChatGPT.
Measuring CSAT before and after
Without a prior baseline, you can't prove improvement. The reasonable measurement:
- Weekly CSAT segmented by channel and ticket category — before rollout, for at least 4 weeks.
- Same measurement for 8 weeks post-rollout.
- Compare with a statistical test (a simple t-test works).
- Qualitative analysis of low-CSAT tickets: what pattern shows up?