Episode Summary

AI agents are becoming remarkably easy to add to customer service workflows, but easy installation does not mean safe deployment. A customer-facing automation that works perfectly during a demo can still hallucinate, access the wrong information, claim to have completed actions it never performed, or trap customers in frustrating conversations with no clear route to a human.

In this episode, Greg Ross-Munro and Justin Davis introduce the FRAME framework, a practical way for service businesses to evaluate AI-powered automation before putting it in front of customers. FRAME stands for Funnel Stage, Risk, Authority, Mechanism of Truth, and Escalation. They explore how each factor affects the amount of autonomy an AI system should receive, why the potential “blast radius” of a mistake matters, how to create reliable connections to systems of record, and why human handoffs, testing, logs, and audit trails need to be designed into the system from the beginning.

Episode notes

  • A frustrating billing experience shows how quickly a poorly designed AI agent can undo an otherwise excellent customer relationship.
  • Automation that helps the customer and automation that primarily helps the business can create very different experiences.
  • AI agents may be easier to deploy safely for intake, scheduling, and pre-sale questions than for complex customer support.
  • Businesses should design clear escalation pathways before launching an AI customer service agent.
  • The FRAME framework evaluates automation using Funnel Stage, Risk, Authority, Mechanism of Truth, and Escalation.
  • The “blast radius” of an AI failure should determine how much engineering, oversight, and autonomy a system receives.
  • Giving an AI read access to business systems carries different risks than allowing it to write, update, or delete information.
  • AI systems can falsely claim they completed actions such as canceling or rescheduling appointments.
  • Confirmation from the actual system of record can provide an additional safeguard for customer-facing transactions.
  • Mechanisms of truth need reliable data, identity matching, integrations, and safeguards against both retrieval and response hallucinations.
  • Customer-facing AI needs uncertainty gating so that it can admit when it does not know or cannot safely answer.
  • Logs, automated tests, audit trails, and periodic human review are essential for finding problems before customers do.

Episode Transcript

Greg Ross-Munro: Hi, Justin.

Justin: Hey, Greg. How you doing?

Greg Ross-Munro: Okay, okay. I’ve got a story to tell you.

I do a lot of judo and jujitsu and things that somebody at my age should not be doing, as you know. So I basically have a loyalty punch card with the orthopedic surgeon and my podiatrist.

My podiatrist is a super nice guy, super cool guy. I had some cracks in my shin and my right foot. There’s a point to this.

I went to him a few times. He did a great job, fixed me up. The last time I went back to him, he looked at me very quickly and said, “You already prepaid your copay at the front desk before the visit?”

I was like, “Yeah.”

He’s like, “You didn’t even need to come. You’re fine. I’ll refund you the copay.”

I’m like, “What? Have you ever heard of a doctor refunding your money?”

He wasn’t angry. He was just like, “You didn’t need to come see me. You’re totally good. I know you had an appointment. Move on with your life. Here’s your money back.”

That’s a great doctor.

A week or two later, I get an email from one of the assistants saying, “You owe us money.”

I’m like, okay, I understand where this is coming from. This will be because of the refund or whatever.

So I send an email back to the automated system. I think it’s automated, or maybe it’s a human. It comes back and says, “If you’re looking for more help, please click here to chat with our live operator.”

I click on that. It takes me to a website. I log in, see a bill that is not mine, and I say, “This is not my bill. I paid this.”

The thing comes back and says, “No, I checked your account records. You must be mistaken.”

It was polite at first.

I’m like, “No. I am 100 percent certain that I not only paid this, but then it was refunded. So what are we talking about?”

No. It is insistent. It says, “You should check your credit card records.”

Justin: It told you to check your credit card records? Hey, dummy.

Greg Ross-Munro: At this point I’m like, is this a human or a bot?

Then I got the receipt, because they gave me a printed receipt, which I kept. I scanned it and sent it back to the thing.

Its response was, “Huh, that’s interesting. This must have come from another company with the same name.”

That was its rationalization.

Justin: It gaslit the hell out of you.

Greg Ross-Munro: Just making stuff up.

This is a service experience that was so good that I walked out of there being like, I’ve never seen a doctor give money back. This guy, if I have any problem ever, if anyone I know has a problem, they’re going to this guy.

But after this experience with the bot, I wanted to immediately jump on Yelp or Google Reviews and write some horrible strongly worded letter because I was so angry.

Does that make me seem insane?

Justin: No. That is an insane experience to have with a chat, to be gaslit into believing, “No, no, no, sir. This receipt was from another company. Bless your heart. You still owe us the money.”

That’s crazy.

Greg Ross-Munro: It just happens to have the same name.

I think there are kind of two kinds of automation out there. There are the kinds that help customers and clients, and then the ones that help you as the business owner at the customer’s expense.

I think clients can tell the difference really quickly.

Justin: Yeah, I think so.

We’re in a new frontier where we now provide automated support via agents and automation, and I don’t know that we have quite caught up yet in terms of understanding the things you have to think about to do that well.

It was only a few years ago when we had a lot of people doing chatbot support via Intercom and Drift with human agents behind them.

Greg Ross-Munro: That was like three years ago, dude.

Justin: Yeah, it might have been three years ago. You’re right.

It feels like the Stone Age now, but it probably wasn’t all that long ago.

Then we sort of said, “Well, there’s this thing that can sound like a human. Let’s put it on the other side of the chatbot and let it do its thing.”

In theory, that sounds great, and in theory it is great. But there are a number of things you have to think about in order to do it well.

If you do it wrong, like this example, it can look very bad for your brand.

Greg Ross-Munro: I agree.

I think it’s probably easier, and there are probably fewer friction points, if you have a service business like a medical spa.

You call the medical spa, an AI answering service answers and books your appointment, and you can ask it questions.

Sorry, this is an example that I saw literally today on a thing being built nearby.

They had a list of all the services that this med spa did, half of which I didn’t understand. But the bot knows what they are and can talk you through the processes, injections, peptides, whatever.

The bot can say, “I can book you in for this,” and it can look at the calendar.

If it doesn’t know something exact, because that’s an entry point, I think there’s almost less friction there than there is on the support side.

The reason is that when I’m calling someplace, maybe I’ve already decided. I’ve already looked. They’re physically close to me. I want that peptide juice inside of me.

I know what I want. Maybe I can ask it some questions like, “Do you have this?”

It’ll be like, “Yeah, we have these peptides.”

You’ve booked me in the calendar. Great. It can send me an email. No problem.

It doesn’t need that much context to do that automation.

But when something goes wrong, that’s when I almost want to speak to a human being.

The self-service stuff outside of an FAQ, or frequently asked questions, is really just a front end to a frequently asked questions system.

Can it actually solve any problems?

Justin: That’s right.

I think that brings us to one of our first principles, or things to think about, when you’re doing customer service with automated systems.

You have to think about the escalation pathways. Think about how those scenarios happen, where they go, and how someone gets to a human in the instance that they need to.

If you’re going to offer human support, ask yourself: At what point do I need to get a human involved? How am I going to know? How is the human going to know? How does the handoff happen?

Just sit down and sketch it out and think about that.

Some AI chat support systems can probably automatically roll back to a call center based on sentiment analysis, where the conversation is going, or the system not being able to answer a question.

Look into that.

But don’t forget that things do need to escalate, and you need to think about what those moments are and how you’re going to handle them.

Greg Ross-Munro: The first part of our framework should be the funnel stage. Where in the process does it live?

That’s going to adjust your thinking around how much automation you use, how much human-in-the-loop you put in, how much engineering you put in, and how sensitive the process is.

If it’s intake or pre-sale, you want speed and clarity. Your failure routes are overpromising, misrouting somebody, estimating, scheduling. These are reasonably well-understood things.

Then I would say the second part of that funnel is the in-service funnel.

You already have a client. You’re sending them a weekly report. You’re trying to reduce friction or prevent rework. You’re trying to prevent manual labor on stuff that’s not really that important but is required to keep the process flowing, like status reports.

Your failure paths there are going to be if there’s a lack of alignment between the actual work being done and what’s being reported.

Then the last one would be the post-service pipeline, which is about protecting reputation and customer or client retention, and deflecting simple work from humans.

My argument would be that the closer you are to money, health, safety, or legal exposure, the less autonomy the bot gets.

That would lead into risk, as we always talk about.

Are you the one who talks about the blast radius of a system?

Justin: I think Connor uses that phrase a lot.

I really like the visual and the metaphor as a way to think about how impactful failure scenarios might be through a risk lens.

For any given system you’re building, it is useful to separate these and think about where they live in the customer lifecycle: pre-customer, purchase, support, service, post-purchase.

But then ask: If an interaction goes wrong, what are the worst-case scenarios that could happen?

We’ll give a couple of simple examples.

Back to the spa example. Not giving all of the details of a service is inconvenient. You might lose a sale here and there, potentially. There’s a little bit of opportunity cost that you could classify as risk.

It’s probably not great, but there’s certainly not much legal risk. It’s a fairly small blast radius.

Now imagine that once you’re a customer, you can talk about your previous treatments. Something causes the system to access customer records in an insecure way.

The AI looks up the wrong John Smith and returns the wrong records of service, outlining embarrassing or otherwise private treatments.

Greg Ross-Munro: Or a HIPAA violation.

Justin: Right.

Now you’ve leaked confidential healthcare information to somebody because of a permissions issue with the system the AI was connecting to.

The blast radius is much larger because now we have legal risk. Somebody could sue you. They would have a right to sue you.

That is a big deal.

Each of these things is not equal. Thinking about the blast radius of a particular automation is very important.

It is so important that it’s worth overdoing to the point where it becomes an intuitive thing that you naturally think about every time you’re considering automation.

Do not leave that one out, because there are some things that are just not worth it for the amount of risk you could incur.

Greg Ross-Munro: When talking about these automations in processes and service companies, we have a framework we like to use for analyzing the risk and how much effort and work should be put into them.

We call it the FRAME framework, which is a little meta.

Justin: A framework.

Greg Ross-Munro: FRAME.

F is for Funnel Stage, so where the automation sits in your process.

R is for Risk, or what we call the blast radius. How dangerous is the thing?

A is for Authority, or how much authority or autonomy the system has.

M is for Mechanism of Truth. What’s the system of integration? What’s the system of record? Where does the truth lie?

Then E is for Escalation, or Error Posture. What happens when it fails? What goes wrong?

Justin: Now that we’ve talked about blast radius and risk, let’s talk about the next piece of the framework, which is Authority, the A.

You can also think of this as autonomy.

What this is about is asking yourself: How much power should I give this system? What can the system do?

One crude way to think about it, and a way we like to think about things in the tech world, is read or write.

Should an automated system be able to read from your other systems, like your EMR or accounting system? Or should it be able to write to those systems, meaning it changes data, potentially deletes things, or updates other things?

Each has downsides and upsides.

Allowing something to read from a system is obviously safer because it can’t hurt the other system in the same way.

Now, it can reveal information that is secure, just like we talked about with our records example.

But the amount of damage it can do in terms of permanently hurting the system or allowing certain kinds of attacks is limited because it can only read from the system.

On the other side, it is also limited in terms of the amount of power, authority, and autonomy it can have in helping solve somebody’s problem.

If somebody needs to reschedule an appointment, for example, just reading from the scheduling system doesn’t help.

The AI needs to be able to write to that system, update the scheduling date, or remove the appointment.

Once you start doing that, you introduce more risk.

Now this thing could schedule it on the wrong day. It could get AM and PM mixed up. It could not represent time zones correctly.

There are these weird things that can happen, and allowing it to have that level of autonomy or authority increases the amount of risk along with the power.

Greg Ross-Munro: There’s another weird thing about this.

Especially if something is driven by a large language model, like an AI chatbot, I’ve noticed that even if they don’t have the authority to do something, they will sometimes, because they’re like eager puppies desperate to help, make you believe that they have done the task.

You want to reschedule that appointment?

“No problem. I got you. Your appointment is rescheduled.”

Your appointment is not rescheduled.

Justin: Right.

Greg Ross-Munro: And if that’s the case, that is terrible for customer service.

Then you show up at the wrong time. Your appointment that you thought was canceled was not actually canceled, which annoys the service provider.

Then the customer shows up and is furious that their appointment isn’t ready because they really wanted that updo or blowout or whatever it is.

Justin: “Updo” is a phrase or word?

Greg Ross-Munro: I don’t know. I try to listen to the things my wife talks about.

Justin: That’s exactly right.

There are a few different things you can do in that scenario.

For something high risk, somebody might say, “I want to move $100,000 from one account to another.”

That is the place where you probably involve a human.

“Great. Let me connect you with an account rep. We’ll get somebody on the phone and help you through that.”

That’s related to escalation paths.

Another thing you can think about doing is a form of what I’m going to call 2FC, not 2FA.

Two-factor confirmation.

Not two-factor authentication.

If you have a scheduling system, your AI bot might say, “Great, I’ve rescheduled you. You will get a confirmation email from that system.”

Then that system should also be responsible for telling the user that they’ve been confirmed, because that system is not going to hallucinate it.

What you probably want is to allow the primary system of record to do the notification and confirmation.

Tell the bot to inform the user to expect a confirmation from that system and, if they don’t see it, to recontact you.

That’s probably the best way to handle that.

Greg Ross-Munro: The thing that’s tough for a lot of people is that there are so many of these systems available off the shelf now.

It’s very tempting to just install one.

You’re running HubSpot in the background and your website is WordPress, and there’s probably a thing that connects those two.

You’re like, “Okay, it’s good enough. I asked it a few questions. Great. This is magic.”

But the QA process and the systems design around it are important.

You cannot trust these things out of the box.

Every business is different. Every business is the same in some ways, and every business is different in other ways.

Those rules and systems are not always going to work out of the box. In fact, they’ll almost never work perfectly out of the box.

Don’t just install one and let it run.

Even if it comes from your billing software or whatever it is, you really need to test and fine-tune it.

Justin: Yes, 100 percent. I could not agree more.

These things are not ready out of the box, and they probably never will be because it would be very difficult for them to be.

There’s always going to be a little bit of fit and finish that has to be done. You could call it last-mile configuration.

It’s an important thing to keep in mind.

Greg Ross-Munro: Let’s move on to the M in our FRAME framework, which is Mechanism of Truth.

We’re talking about what the actual system of record is and how strong the integration with it is.

Do you have any thoughts on what could be useful for people to know about how to treat your source of truth, or your mechanism of truth?

Justin: The biggest thing is that you have to watch out for a few things.

We’ve talked about hallucinations before, but we’ll go over it again.

Hallucinations can come in a couple different types. What I would call pre-retrieval and post-retrieval hallucinations.

A pre-retrieval hallucination is where the system queries the mechanism of truth incorrectly due to poor prompting, poor context design, or something like that, and gets the wrong user record or wrong record back.

The fetch is incorrect.

A post-retrieval hallucination is where, after it has made the call to the database, whether the call succeeded or not, what the LLM does on return becomes important.

Did it get the right data? How can it validate that it had the right data?

If it doesn’t, because of sycophancy issues it could hallucinate and make up data.

It could tell you things that are not true because maybe the set came back null and it is trying to make sure it always returns some data.

You have to watch out for those two types of hallucinations when we’re talking about mechanism of truth.

Greg Ross-Munro: Is that something you handle mostly in the prompt?

You tell the system, “If you don’t know the answer, or if null is returned, say that.”

Or is it an integration problem where the connection to the system is bad or maybe the underlying data is bad?

Justin: I think it’s all of that.

Fundamentally, the thing about LLMs is they’re what we call nondeterministic, which means they won’t always return the same thing.

If you ask it twice, it won’t necessarily return the same thing every time.

What I think you want to do with those kinds of nondeterministic systems is build as much determinism into them as you can, as close to the LLM layer as you can.

What do I mean by that?

Maybe the software you’re using allows you a way to easily connect with a system, provision what you can access, and do all of that in a structured way.

If it has real first-class support for those integrations, that’s going to be more trustworthy.

There’s a little bit more focus on the prompt and the pre-retrieval part of this.

In other examples, if you’re rolling your own connection to something, instead of doing all of it in a prompt, build it as a skill or build it as an MCP server.

You’re introducing some kind of determinism into the stack so the LLM is making fewer decisions when using the toolset offered to it.

That gives you a more reliable chain to and from the mechanism of truth.

Greg Ross-Munro: I think you can make some decisions based on that.

You could grade the data you have and the integrations you have and work out what you should do.

Give it an A grade if the data is fresh, it’s complete, and you have a reliable way to match the identity of the user in your system. Then you can make more definitive statements.

If you give it a B, maybe there are some minor delays in updating the data or some partial fields missing. That’s when you have to be cautious about the claims you’re making.

If you give it a C or lower, you’ve got stale data, incomplete data, or fuzzy identity matching.

Then you’ve got to prompt the thing to avoid making definitive claims.

If it’s about anything important, like eligibility, disputes, or account balances, you gate the confidence of those responses based on the grade you gave the integration and the reliability of the identity match.

Then the last thing is, what happens when something goes wrong?

You’ve got escalation or errors.

What happens when things go wrong? How do you stop things from going wrong?

Justin: That’s the E in the FRAME framework: Escalation and Error Posture.

What you have to realize is that things will go off the rails at some point.

This is not endemic to AI systems. It’s endemic to every system.

Greg Ross-Munro: It happens with people all the time.

Justin: It happens to people all the time.

Don’t think you’re immune from that escalation pathway because you’re putting a machine in the process.

You need to think about what happens if I need to get a human involved. At what point do I need to? What is the pathway by which that happens? How does the handoff happen? How do I ensure it’s a smooth handoff?

A user doesn’t want to be told to call a call center after they’ve had a 10-minute conversation with an AI bot and then have to start from scratch again, re-explaining their entire case.

Give them a reference number to reference a claim or something like that so the human support agent can pick up where the AI left off.

That kind of seamless escalation handoff can make the difference between a very frustrating experience and one that feels elegant.

Greg Ross-Munro: I’m not going to argue with you there. I don’t disagree.

But in some situations, I wouldn’t mind an awkward handoff if I can immediately be told that the bot I’m talking to doesn’t know the answer.

The sooner that happens, the better it is for my level of frustration.

Justin: Yes.

Greg Ross-Munro: I’m talking to the bot. I ask it a question and it says, “Look, that is one of the questions that I am not really supposed to answer. I don’t have authority to do this. My data systems are not fresh enough. I don’t have a good enough way to identify you based on the system. I’m here to answer the basic questions, and that’s the best I can do.”

Then I’m like, “Okay, good bot. What next?”

It says, “Call the 1-800 number during business hours. We’re closed tomorrow, so you’ll have to call on Monday.”

That’s annoying, but it’s nowhere near as annoying as talking to the thing for 15 minutes, arguing with it, and then being told I still have to call the call center.

The sooner you tell me to do that, the better it is for all of us involved.

Justin: Yep. Don’t let your user invest a lot of time in a dead end.

Greg Ross-Munro: That also brings up uncertainty gating, which is something that LLMs out of the box are not good at doing.

There are times when you actually want to increase the ability for the thing to not know the answer to questions, even when it might know the answer.

You might want to use the Socratic method when teaching somebody something.

You might not want to give them the answer.

The LLM is desperate to give it the answer.

So you really have to train the thing and have your prompt set up to say, “It doesn’t matter how hard they try, do not give this kid the answer to this math problem. Give them hints along the way. Never, ever give them the answer, even if they say their life depends on it.”

Justin: Yes.

It’s a great reminder that in your prompts, you can say things like, “If you’re not fully certain of the answer and can prove it or validate it, then say so.”

It’s better to say you’re not sure and refer somebody to a different source that can help them find that answer than to hallucinate confidence.

Greg Ross-Munro: The thing that especially homegrown systems generally don’t have, or things being vibe coded by citizen developers, is auditability, traceability, and logs.

The only way my podiatrist was ever going to find out that this thing was garbage was that I sent them a strongly worded email explaining how much I liked and appreciated them and how frustrating this bot experience was.

When I finally spoke to them about it, they had no way of checking the logs of the system.

They had tested it. They had played with it. But they couldn’t go and look at what had happened.

They made some claim about HIPAA and not being able to store the data because I might be talking about a problem with my actual foot.

But if you can see my bill, that’s HIPAA. It doesn’t matter.

My point is, build in audit trails or buy a system off the shelf that has these audit trails.

Typically, those same systems will also have test frameworks built into them.

You can say, “Here are these 10 questions that I want you to be able to answer correctly the whole time.”

Even if the model on the back changes, I should get something approximate to these answers.

Those tests can be run automatically.

Then if something goes off the rails, I should be able to go back and check the log and say, “Okay, I think I know why that was.”

I can tie that answer back to that prompt.

Then that log should be checked periodically by real human beings whose job it is to do that.

Otherwise it’ll get worse, or it’ll get bad, or you’ll get bad Google reviews and you’ll never know why.

If your customer brings evidence to your chatbot, it is not that bot’s job to debate it.

It is time to start routing.

Justin: Yes.

Hopefully the FRAME framework has been helpful.

Just to restate it:

F is Funnel Stage. Where is the customer in their journey, and what type of task in the funnel is the system helping them with?

R is Risk, the blast radius. How bad is it if this thing gets something wrong in this interaction?

A is Authority. How much authority and autonomy do you give this thing to act on your systems on the customer’s behalf?

M is Mechanism of Truth. What is the system of record? How is it integrated? How do you protect that integration? How do you verify that it’s doing what it’s supposed to be doing?

And finally, E is Escalation. What is the error path when these things go wrong and you need to escalate to a human? How are you going to do it? What’s the handoff look like?

How do you make sure you can audit it afterward to make sure that you’re correcting and improving the systems on an ongoing basis?

This stuff takes real thought. It takes real engineering. It takes real design.

These are not necessarily easy, simple things to do.

With small things, it can be simple. But the bigger and more business-critical they become, the more thought this takes.

Don’t just grab something off the shelf and think, “I’ll launch this tonight.”

Take a minute.

Greg Ross-Munro: Don’t YOLO your AI bot.

All right. We’ll post the framework on the blog and show notes, and hopefully it’s useful to somebody out there.

Justin: Excellent.

Greg Ross-Munro: All right, Mr. Davis. It’s good to see you, sir. Have yourself a great weekend.

Justin: Good to see you. Enjoy your weekend. I hope it’s filled with fighting.

Greg Ross-Munro: There will be some fights. Hopefully my foot holds up. Otherwise I’ll just go argue with the bot.

Peace.

Justin: Excellent. All right. We’ll see you guys.

Greg Ross-Munro: Ciao.

Comments are closed.