Pitcocy
ai-agentsai-systemse-commerceemail-automation

AI Email Agent for Ecommerce: What Stops a Wrong Reply

ยท14 min read

๐Ÿค–Explore this article with AI

Get AI-powered summaries, insights, and analysis from top AI platforms.

Every webshop owner I talk to about AI and support email has the same fear, that their AI agent will write something stupid to clients and it will cost money and even worse REPUTATION, which you can't build back up anymore.

Stuff like a refund promised that nobody agreed to, order details sent to the wrong person, a cheerful answer to someone who just mentioned their lawyer.

That fear is fair, and "the model is really good now" is not an answer to it.

I built an AI email agent for a webshop that sells trailer parts in 26 countries, and most of the work was not getting it to write good replies. Most of the work was everything that sits between the agent and the Send button. This post walks through those layers, so you know what to ask for before you let any AI near your support inbox.

AI email assistant or AI email agent?

But first what is the key difference between these 2:

An AI email assistant helps you with your own inbox. It summarizes threads, suggests a reply, sorts your newsletters. If it gets something wrong, you notice, because it's your mail.

An AI email agent works a shared support inbox. It reads the customer's message, looks up the order, checks your policies, writes the reply in the customer's language and, if you let it, sends it. It acts for your business, which is exactly why it needs rules that an assistant doesn't.

This post is about the second one.

The shop: 26 webshops, one support inbox

TrailerPlus sells trailer parts: axles, brakes, couplings, wheel bearings, basically everything that makes your trailer move.
They run 26 country webshops and the agent supports 28 languages.

Dennis runs support, and the emails come in from all those countries, each in the customer's own language.

The agent now reads each incoming email, sorts it into one of six categories, gives it a risk colour and writes a draft. The full build is on the TrailerPlus project page, here I want to stay on the safety side.

The screenshots in this post come from a demo copy of the system with made-up customers.

Human in the loop AI: every reply gets approved, in Dutch

"Human in the loop" sounds good on a slide. In practice it only works if the human can read what he is approving.

A customer writes in Portuguese, the agent drafts in Portuguese, and Dennis doesn't read Portuguese. Nobody reads 28 languages. If the review screen shows him a draft he can't read, he is not reviewing anything.

So the whole review side runs in Dutch. Dennis sees the customer's message and the agent's draft side by side, both in Dutch. When he wants to change something he edits the Dutch text, the system translates his edit back into the customer's language, and the Send button stays disabled until that translation is done.

This is the part I'd look at first in any tool you consider. If you sell in more than one language, the review screen has to work in yours.

It also gave me my favourite bug of the project. A draft for a customer in Portugal correctly listed the payment methods for Portugal. The Dutch version that Dennis reads said the customer could pay with iDEAL, which is how the Dutch pay. The translator was being helpful and had adapted the facts for a Dutch reader :D

The reply to the customer was fine, the copy for the reviewer was wrong, and a reviewer who reads the wrong thing makes the wrong call. The translation step now has one hard rule: translate, don't localize.

Six layers between the AI customer service agent and your customer

Customer support automation goes wrong when the only thing between the model and your customer is the model's own judgment. So I stopped asking "is the model good enough?" and started asking "what is the worst email this could send, and what catches it?"

Every layer below can only do one thing, and that's push a reply back to a human. None of them can make a reply go out.

"What if a customer threatens with a lawyer?"

There's a list of words that always force a human review. Legal threats, chargebacks, journalists and social media, consumer watchdogs, fraud accusations. The list exists in 15 languages, and when you add a new term the system translates it into the others for you to check.

The matching is plain text and there is no AI involved at that moment. That's deliberate, because you can talk a model out of things and you can't talk a word list out of anything.

"What if it promises a refund nobody agreed to?"

Word lists don't catch everything. A customer can be furious without using a single word on the list, and the agent can write a friendly reply that quietly commits you to something.

So after all the hard checks pass, a second and separate AI call reads the customer's message next to the proposed reply. It looks for anger the first one played down, for promises of refunds or replacements, and for questions the reply skipped. It has to write down its reason in one sentence so Dennis can see why a draft was held.

And if that check itself fails or times out, the reply is held. A broken safety check counts as a no.

"What if it sends order details to the wrong person?"

Anyone can email a shop with an order number they found or guessed and ask where it's going.

The agent can look up orders and trigger an invoice resend, but both sit behind an identity check that runs on the server. The order has to belong to the email address the message came from. The model doesn't get to choose whose data it sees, and it can't be talked into checking a different address, because that part is not up to the model at all.

When the check fails, the agent gets nothing about the order and asks the customer to write from the address the order was placed with.

The decision I like most here is what happens when the shop's order system is down. The agent doesn't guess and doesn't share what it half knows. It says less, tells the customer a colleague will follow up, and the draft is held. The same goes for product links, which are built by the app from the right country webshop and never typed by the model.

"What if it answers a stranger nobody ever looked at?"

The agent never sends anything on its own to someone a human hasn't replied to before. First contact always gets a person.

What counts as trust? The AI already gives every email a risk colour and a confidence score, so the easy route is to let a "green, very confident" email through. I didn't want the model to decide who the model gets to email. Only a reply that a human reviewed and sent makes a sender trusted. A thumbs up on a draft doesn't count, and neither would a "trust this sender" button, because it would skip the one thing that matters, which is a person reading a real reply before it goes out.

"What if it goes wrong at 3am?"

The classic one is two robots talking to each other. Your agent replies, the customer's out-of-office replies back, your agent replies to that, and by breakfast there are 200 emails in the thread.

Three things sit here. There are caps on how many replies can go out on their own per hour, per day and per customer. There's a loop detector that spots auto-replies and answers that come back faster than a person can type, and locks that thread until a human releases it. And there's a kill switch that sends everything back to draft with one click, while Dennis can still send by hand.

If this sounds familiar, it's the same idea I use for Google Ads agents, where the agent can propose a budget change but has no button to apply it. That one is on the PPCOS project page.

You decide how much the AI agent for ecommerce does alone

The sixth layer is you.

You wouldn't hand a new hire the keys on day one and leave for a holiday. You'd read their first emails, then stop reading the easy ones, then only look at the tricky ones. The system is built to work the same way.

It has three modes. Off means nothing leaves the building, not even when you click Send. Shadow means the agent drafts everything and a person sends everything. On means replies that pass every layer above can go out on their own.

Even in "on", nothing goes out by default. Each of the six categories has its own switch and every switch starts off. You turn them on one at a time, starting with the boring ones.

At TrailerPlus, delivery and shipping questions are the first category picked to go alone. The system is still in shadow today, so a person still clicks Send on every reply. Dennis flips it when he trusts it and not when I say so.

How do you know when to trust it? You look at how often you send the draft without touching it. In the month we measured, about four in ten drafts went out exactly as the agent wrote them, and the rest got edited first. That number tells you more than any demo.

It should also go up over time. Dennis asked whether the agent could learn from the answers he gives and edits, so that's the newest part of the build. Every reply he sends is stored as an example, with names and order numbers stripped out. When a similar question comes in, the agent sees how Dennis answered it last time. It went live this month, so I have no before and after to show you yet.

The agent doesn't learn from replies it sent on its own. It only learns from replies a person reviewed, otherwise it would be copying its own homework.

This is the same ladder I describe in how to build an AI system for Google Ads: observe first, then recommend, and only then execute.

When off-the-shelf customer support automation is enough

I build custom systems, so take this with that in mind, but I'm not going to tell you everybody needs one.

If you can plug a tool into your helpdesk, change two settings and it does the job, use it. It's cheaper, it's live this week, and somebody else maintains it.

The question to ask is how much of the above you need, and how much of it you'd have to build around the tool yourself. Can you review in your own language? Can it check that an order belongs to the sender in your shop system? Can you switch autonomy on per category, or is it all or nothing? Can you see why a reply was held?

If the list of "we'd need to work around that" gets long, you're building a custom system anyway, just in somebody else's settings screen.

The other thing a custom build gives you is independence. When it's done you get the codebase and I show you how to change it. You can run it and manage it however you want, with me or without me.

Questions webshop owners ask me

When does a webshop need an AI email agent?

When the same kinds of questions come in every day and answering them means looking things up: order status, delivery times, which part fits. Add more than one language and it gets worth it quickly. If you get five emails a day, you don't need this.

Can an AI email agent send emails on its own, and do I have to let it?

It can, and you don't have to. You can run it in draft mode forever and still never start a reply from a blank page. Sending on its own is a switch per category that you turn on when you're ready.

What happens if the agent gets a reply wrong?

In draft mode you catch it in review, fix it and send your version, and your fix becomes an example for next time. For replies it would send on its own, that's what the layers are for, and every decision to send or hold is logged with the reason so you can check afterwards.

Will it replace my support team?

  1. Dennis still decides what goes out. What changed is that he reviews and edits drafts in his own language instead of writing every reply from zero in someone else's. The angry customers, the edge cases and the judgment calls still land on a person, and they should.

What if it doesn't work for my shop, and what does it cost?

That's what a pilot is for. We pick one workflow, I build it end to end on your real email, and you get a fixed quote once the workflow is mapped. If the pilot doesn't run in production, you don't pay the final invoice.

Am I locked in?

  1. You get the code and the guidance to change it. Categories, tripwires, the knowledge base and the tone instructions are all editable from the dashboard without a developer.

Can ChatGPT reply to customer emails?

It can write a decent reply if you paste the email in. It doesn't know the order, it doesn't check who is asking, and the only thing between its answer and your customer is you with copy and paste. For a handful of emails a day that's fine.

What is the difference between an AI email assistant and an AI email agent?

An assistant helps one person with their own inbox. An agent works a shared support inbox, uses your shop data and can act for your business. The second one needs the layers in this post.

Start with one category

If the fear of a wrong reply is what's keeping you from trying this, good, keep it. Just turn it into a list of questions: who reads the reply before it goes out, in what language, what gets held, and who decides when that changes?

If you want that built around your shop, that's what the pilot build is for. One workflow, on your real inbox, and if the pilot doesn't run in production, you don't pay the final invoice :D

Alfred Simon

About Alfred Simon

AI Systems Builder & Coach

I build custom AI systems for companies: support email, order processing, content, reporting. I write about context management, AI workflows, and the messy reality of building things with AI. No theory. No hype. Just what survived 30+ agents and a very healthy trash pile :D

Want to build something like this for your team? Let's talk.

Want AI systems that hold up in production?

Whether you run a team or work solo, I can help you make AI useful for your work.