An AI agent can be technically capable of clicking a button and still be the wrong system to trust with that button. Most automation demos skip this part. Building our own Sales OS made it impossible to skip.
Here are the four failures we designed against, plus one we found in our own setup. None of the fixes are clever. That’s the point.
1. Invented research
What it looks like: the model writes a warm, specific compliment (“loved your talk on pricing last month”) that no source supports. It reads as personal. It’s fiction.
Why it matters: one invented detail in a cold message destroys trust faster than a generic one. It can also be simply wrong about a real person or business.
The guardrail: research is saved with an evidence URL. Drafts are generated only from saved research, and every draft shows the evidence it used. If the evidence is weak, the system doesn’t manufacture a message. It flags the lead instead.
2. Duplicate delivery
What it looks like: two background jobs pick up the same approved message at the same moment. The prospect gets it twice. Or a retry fires after a send that actually succeeded.
Why it matters: a duplicate looks careless at best and like spam at worst.
The guardrail: a send lock in the database. A message is claimed for sending in a single atomic step, so only one job can send it, and a retry checks that lock before trying again.
3. Secret exposure
What it looks like: an API key ends up in browser code, a public repository or a screenshot in a demo video.
Why it matters: an exposed email or database key lets someone else send as you or read your data.
The guardrail: secrets live only in server-side environment variables. The browser never receives a service key. Demo recordings use fictional data and hide keys, tokens and database identifiers.
4. Uncontrolled retries
What it looks like: a job fails, retries, fails again. Forever. Each attempt sends a little more, or hammers an API until the account is blocked.
Why it matters: runaway retries turn one error into spam, cost or a suspended account.
The guardrail: a retry limit and a dead-letter state. After a fixed number of attempts, a job stops and waits for a person. Daily sending caps sit on top of that, plus a dry-run mode for testing and an emergency pause that stops all sending at once.
The one we found in our own setup
We had an older scheduled task designed to connect, follow and message people on social platforms automatically, with nobody watching. Even with approved message text, that removes the final human judgement at the moment of sending, and it breaks the platforms’ rules on automation.
We disabled it. Social outreach in our system is now manual: the system prepares the draft, a person opens the profile, sends the message and marks it done.
The line we draw
The rule that ties all of this together is simple: AI can prepare. A person must approve anything consequential.
| AI can prepare | A person must approve |
|---|---|
| Research summaries, with sources | Every outbound message |
| Fit scores and qualification notes | Any social-platform action |
| Draft emails and follow-ups | Changes to permissions or data access |
| Reminders and next-step suggestions | Anything involving billing or money |
This makes the system less impressive in a demo. It also makes it something a real business can trust with its name.
AI is powerful. But a powerful system without controls isn’t a business asset. It’s a business risk.
Want to see these controls working? Watch the Sales OS demo, or read how we built it. If you’re planning automation in your own business, tell us about the workflow first.