Engineering & Implementation 6 min read
What makes an AI system reliable?
When you put an AI assistant in front of customers or staff, it speaks for you. Its mistakes are your mistakes.
In short
When Air Canada's chatbot misstated bereavement travel rules, a Canadian tribunal found the airline responsible. What separates an AI assistant you can trust from one that quietly makes things up?
The Air Canada ruling
In February 2024, a Canadian tribunal ordered Air Canada to compensate a customer because of something its website chatbot said.
The customer had asked the chatbot about bereavement travel after a death in the family. The chatbot gave an answer that did not match the airline's policy. Air Canada argued the chatbot was responsible for its own words. The tribunal disagreed: the chatbot was part of the airline's website, and customers shouldn't have to double-check one part of a company's website against another.
The lesson is clear. When you put an AI assistant in front of customers, staff or beneficiaries, it speaks for you. Its mistakes are your mistakes.
So what separates an assistant you can trust from one that quietly makes things up?
What failure looks like in practice
AI failures rarely look dramatic. They look like this:
- A confident answer that is simply wrong
- An answer drawn from general internet knowledge instead of your own policies
- A follow-up question that gets confused with the one before it
- An answer that shows someone information they were never meant to see
- An "I can help with that" when the honest answer is "please speak to our team"
Field notes from our own testing
We recently tested the AI assistants on several clinic websites, using simple questions a real visitor would ask. On general health topics the answers were generally sound. But the assistants often struggled with the clinic itself. Asked for the clinic's phone number, one gave a national helpline instead, while the clinic's own number was on screen. Others couldn't give weekend opening hours, or said a service wasn't offered when the clinic had a page describing it.
None of this is unusual. The assistants were trained to sound knowledgeable, but nobody had grounded them in the one source that mattered most: the clinic's own information.
The five pillars of a reliable AI system
Reliability is an engineering discipline built upon five fundamental pillars:
- 1. Grounding: Answer from your sources, and show them. A reliable assistant answers from your documents, website, policies and data, not from whatever the underlying model happens to know. It should be able to show where each answer came from, so a person can check it in seconds.
- 2. Retrieval quality: Find the right passage first. Most business assistants work by searching your content for the relevant passage and then writing an answer from it. If the search finds the wrong passage, even a very capable model will produce a wrong answer. In our experience, many 'the AI is hallucinating' problems turn out to be search problems.
- 3. Knowing when to stop: The most important thing an assistant can say is 'I don't know, let me connect you with someone who does.' A reliable system has clear limits: topics it won't answer, confidence thresholds below which it hands off, and a clean route to a human through a phone number, email or ticket.
- 4. Access control: The right answers for the right people. Inside an organization, not everyone should see everything. HR files, donor records, beneficiary data and board papers need protecting. A reliable internal assistant only draws on documents the person asking is cleared to see. That is designed in from the start, not added later.
- 5. Evaluation: Test with real questions, and keep testing. A demo with five friendly questions proves very little. Before launch, a reliable system is tested against a large set of real, messy questions, including ones it should refuse and ones it's likely to get wrong. Every change to the content, the model or the settings is tested again. If you can't measure how often it's right, you don't know whether it's reliable.
"But doesn't connecting it to our documents fix this?"
It helps a great deal. It doesn't fix everything.
In 2024, Stanford researchers tested leading AI legal research tools, built by major legal publishers and grounded in their own case-law databases. Some had been marketed as avoiding or even eliminating made-up answers. The researchers found that each tool still produced incorrect or unsupported answers between 17% and 33% of the time on their test questions. They did better than a general-purpose chatbot, but they were far from perfect.
Grounding reduces errors. Good search, clear limits, human handoff and ongoing testing are what bring them down to a level you can live with.
Five questions to ask any AI vendor
Whether you're buying a tool or commissioning a custom build, ask:
- 1. Where do the answers come from, and can I see the source for each one?
- 2. What happens when the system doesn't know the answer?
- 3. How was it tested before launch, with how many questions, and what was the error rate?
- 4. Who can see what? How is sensitive information kept from the wrong people?
- 5. How will we know if quality drops after launch?
If you already have an assistant running
Try this today: Write down ten questions your customers or staff actually ask, including two or three awkward ones. Put them to your assistant and check each answer against your own records. Note any answer that is wrong, invented, or should have been passed to a person.
If more than one or two fail, the problem is very likely fixable, but it needs fixing before it costs you a customer's trust.
The takeaway
Reliability doesn't come from choosing the cleverest model. It comes from design: grounding answers in your own information, finding the right passage, knowing when to stop, protecting what should be private, and measuring the outcome rather than the demo.
Already using an AI assistant and not sure you can trust it? Our AI System Health Check tests retrieval, grounding, accuracy and efficiency, then gives you a clear findings report with a prioritized fix list. Write to us at ada@blueberryia.com.
Sources: Moffatt v. Air Canada, 2024 BCCRT 149 (Civil Resolution Tribunal of British Columbia, February 2024); Magesh et al., 'Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools,' Stanford University (2024).