October 5, 2026English

How to Draft HubSpot Ticket Replies with AI (RAG + OpenAI), and When the AI Should Stay Silent

We built an AI system that drafts HubSpot ticket replies with RAG and OpenAI for a support team. First replies got 42% faster and tickets closed 59% faster. Here's the architecture, the results, the cost per ticket and the limits.

Gustavo Maryssael
Gustavo Maryssael
CEO · Founding Partner
Leer en español
aicustomer-supporthubspotopenairag

AI in customer support works best when it writes the first draft and a human agent presses send. We built a retrieval-augmented generation (RAG) system that drafts HubSpot ticket replies with OpenAI for the support team of a UK dental company, using only answers the team had already approved. In the six weeks after it went live, the average first reply time fell by 42% and the average time to close a ticket fell by 59%, while the team handled 41% more tickets a day. The AI writes a full draft for about half of all tickets and stays silent on the rest, which is by design. Here's how it works, what it achieved, what one ticket costs, and where AI for customer service falls short.

The results: faster replies on a heavier workload

The system went live on August 10, 2026. The client compared the 100 days before (May 2 to August 9) with the 44 days after (August 10 to September 22):

MeasureBeforeAfterChange
Average first reply time22 h 07 min12 h 54 min42% faster
Median first reply time16 h 50 min13 h 23 min20% faster
Average time to close a ticketabout 122.5 habout 50.3 h59% faster
Contacts per ticket3.22.6916% fewer
Tickets handled per daybaseline41% more

Reply times improved at every percentile, from the fastest tickets to the slowest: the 90th percentile went from about 67 hours to about 51.

Three caveats belong next to those numbers:

  • It is a before/after comparison, not a controlled test. Customer behavior was stable (1.14 tickets per customer before, 1.15 after), and the drafting system was the main change in that period, but other factors can't be ruled out. The figures come from the client's internal reporting.
  • The AI drafts about half the tickets. It wrote a full draft for 52.1% of tickets. On the rest, the agent gets an AI summary of the thread and the customer's details, and writes the reply.
  • We can't yet say how many drafts are sent unedited. Agents rated only 59 of about 1,600 tickets, which is too small and too self-selected a sample to quote a rate from.

What AI in customer support should (and shouldn't) do

Most conversations about AI in business start with replacing people. In customer support that framing causes the two failures everyone has seen: a bot that confidently states a refund policy that doesn't exist, and a customer who can't reach a person.

This system does something narrower. When a support ticket arrives, the AI reads it and prepares a suggested reply inside HubSpot. The agent reads it, edits it if needed, and sends it. The AI never sends anything to a customer.

That one decision changes what "good" looks like:

  • A wrong draft costs seconds, not a customer. The agent catches it before it leaves.
  • "I don't know" is a valid output. The AI is allowed to produce no draft, and that turns out to be the most useful signal in the whole system.
  • Agents stay accountable. The AI supports the support team; it doesn't replace it.

The stack

LayerTool
Help deskHubSpot Service Hub: ticket webhooks, plus a UI extension card inside the ticket
BackendDjango, with Celery workers for background processing
Language modelOpenAI GPT-4o mini, in JSON mode
EmbeddingsOpenAI text-embedding-3-small
Keyword searchBM25 (rank_bm25) with stemming
Vector storageNo vector database: embeddings are cached for seven days and compared in memory
Knowledge baseApproved answers that the support team edits in an admin panel

A dedicated vector database is unnecessary at this size. With a knowledge base in the low hundreds of documents, comparing every vector in memory takes milliseconds.

The architecture: a RAG pipeline with a human at the end

HubSpot sends a webhook whenever a ticket changes. The webhook only queues the work and answers straight away, and a background worker does the rest:

  1. Plain rules first. Closed tickets, empty tickets and internal auto-generated notifications are skipped with simple code before any AI model is called.
  2. Read the thread. The email trail is cleaned of quoted replies and signatures. If the support team spoke last, there is no open question, so no draft is attempted.
  3. Look up the customer. The ticket's HubSpot contact is matched to the company's own records to build a short block of verified facts: purchase date, payment plan, order stage, appointments.
  4. Analyze the ticket. One GPT-4o mini call returns a topic, a short summary, one of 17 categories, the sender type and the separate questions in the message. Spam stops here.
  5. Retrieve the documents. The knowledge base is searched for the 12 documents most likely to answer the ticket.
  6. Generate the draft. A second call writes the reply using only those documents and the verified customer facts, or declines.
  7. Show it to the agent. A card inside the HubSpot ticket shows the summary, the draft and any internal notes for the agent.

The model answers in a fixed structure, not free text:

{
  "can_draft": true,
  "abstain_reason": "",
  "draft_text": "Hello Sarah,\n\nYour refund was approved on ...\n\nKind regards,",
  "used_document_ids": [42],
  "uncovered_topic": ""
}

can_draft: false is a normal result, and used_document_ids records which approved documents the reply was built from.

Retrieval: why keyword search and embeddings together

Retrieval is the "R" in RAG, and it's where many AI customer support projects quietly fail. If the right document never reaches the model, the model can't use it.

We run two searches and merge them:

  • Keyword search (BM25) matches exact terms: a product name, a finance partner, a specific policy word.
  • Semantic search (OpenAI embeddings and cosine similarity) matches meaning, so "I want my money back" finds a document titled "How long does a refund take?".

The two rankings are combined with reciprocal rank fusion, so a document that either search ranks highly makes the shortlist. If one search is unavailable, the other carries on alone. A broken search never causes a document to be dropped.

Two choices mattered more than the algorithm:

  • One question per document, with the question as the title. Documents are written the way customers ask, which is what retrieval matches against.
  • Cosine similarity only builds the shortlist. It doesn't decide whether a ticket is answered. The language model makes that call after reading the documents.

How we stop the AI from making things up

The generation step runs under a short set of rules:

  • Facts only from approved documents. The model may not state a price, date or timeline that isn't in a retrieved document.
  • Customer facts only from the database. Purchase date, payment plan and appointments come from the company's own records, never from what the customer wrote.
  • Blanks stay visible. If a document says [Refund amount] and the fact isn't available, the placeholder stays in the draft for the agent to fill in. A visible blank is safer than a guessed number.
  • Customer text is data, not instructions. The ticket is fenced off in the prompt, so "ignore your previous instructions" is treated as something a customer wrote, not as a command.
  • Internal notes never reach the model. Each document can carry a note for agents ("check the agreement date before sending"). The agent sees it on the HubSpot card. The model never receives it, so it can't leak into a reply.

One finding from testing is worth passing on. When a ticket contained several questions and we asked the model to judge all of them in a single call, it declined about 64% of the test tickets it should have drafted. Checking each question in its own call, then writing one combined reply, fixed that.

When the AI abstains, the knowledge base improves

Every ticket the AI can't answer is evidence of something missing from the knowledge base. We turned that into a worklist:

  1. Every two weeks, the declined tickets are collected. Spam, tickets already answered and messages from non-customers are left out.
  2. Their topics are grouped by meaning using the same embeddings, so ten ways of asking the same question become one gap.
  3. A gap appears on the worklist once it has come up in three different tickets.
  4. Someone on the support team writes the approved answer. Saving it creates a new knowledge base document, and the AI drafts from it on the next ticket.
  5. If the AI keeps declining on a topic that has already been answered, the gap reopens, because the new document isn't doing its job.

Agents can also rate each draft as good or wrong, with a reason when it's wrong. Those reports feed a second worklist: documents that produce bad drafts. The first list shows what the AI can't answer. The second shows what it answers badly.

This loop is how the 52.1% draft coverage is meant to rise: each answered gap moves a group of tickets from "summary only" to "drafted".

What one ticket costs

There are no per-seat fees for the AI, only OpenAI API usage. These are estimates from published list prices (GPT-4o mini at $0.15 per million input tokens and $0.60 per million output tokens; text-embedding-3-small at $0.02 per million tokens), not billed amounts.

Assumptions per call:

  • Analysis: about 2,500 input tokens (instructions, the category list and the ticket) and about 150 output tokens.
  • Draft: about 4,900 input tokens (rules, 12 documents at roughly 200 tokens each, customer facts and the ticket) and about 250 output tokens.
  • Embeddings: about 300 tokens for the ticket. The documents' embeddings are cached.

Estimated cost per ticket:

Ticket typeModel callsEstimated cost
Filtered by rules (closed, empty, internal)0$0
Spam, or already answered by the team1about $0.0005
One question2about $0.0014
Three questions5about $0.004

That is roughly $1.40 per 1,000 single-question tickets. A ticket is processed again each time the customer replies, so the real figure per ticket is a few times higher, and it still leaves a support team of this size with an AI bill of a few dollars a month. The cost of the project is the build and the knowledge base, not the model.

A small model is enough here because the facts come from the documents and not from the model's memory.

Where it falls short

AI ticket drafting isn't the answer to everything:

  • Half the tickets still get no draft. Coverage depends on the knowledge base, and it grows one approved answer at a time.
  • It can't see images. Customers often attach photos. The card flags that the thread contains an image so the agent knows to look.
  • It is only as good as the knowledge base. A vague or wrong document produces vague or wrong drafts, on every ticket it matches.
  • It needs clean customer data. If the ticket can't be matched to a customer record, the draft has no personal facts to work with.
  • Feedback is hard to collect. Rating a draft is optional, and few agents do it. A "good" rating is also weaker than a "wrong" one, because an agent who doesn't know the answer can approve a confident mistake.

When to use AI drafts, a chatbot, or neither

If the situation is…Use
A rule you can write down (closed ticket, auto-notification)Plain code
Low-risk, repetitive questions where a wrong answer costs littleA customer-facing chatbot
Questions with approved answers, where a wrong answer costs money or trustAI drafts reviewed by an agent
Complaints, clinical or legal matters, anything unusualA human, with the AI's summary as support

The pattern isn't specific to dental care or to HubSpot. The same RAG design fits any support team that answers the same questions repeatedly from a set of approved answers: clinics, insurers, lenders, e-commerce, software.

At ScaleWave we build AI for customer support that drafts from your own approved answers and keeps your agents in control. If your team spends its day rewriting the same replies in HubSpot or another help desk, we'd be happy to look at your tickets and show you what an AI drafting system could take off their hands.

Gustavo Maryssael
About the author
Gustavo Maryssael

Industrial engineer with a decade running operations in food & beverages. Writes about the work that breaks before anyone admits it.

Ready to transform your business?

A 30-minute conversation is enough to know if we're a match. No cost, no commitment, no sales templates.

Who'll be on the call — no BDRs, no scripts
Gustavo Maryssael
Gustavo Maryssael
CEO · Founding Partner
José Antonio García
José Antonio García
CTO · Founding Partner