AI capability is not a smooth ladder. It is a jagged frontier.

A system can perform exceptionally well on one task and fail unexpectedly on another that appears equally difficult to a human. It may write a strong summary of a supplied report but invent a fact when asked an open-ended question. It may classify a familiar document correctly but misread a rare exception. It may solve a short workflow but lose coherence in a longer sequence of dependent steps.

This uneven boundary is one of the most important ideas for enterprise deployment.

Why averages mislead

Teams often evaluate an AI system through averages:

  • Average accuracy
  • Average time saved
  • Average user satisfaction
  • Average benchmark performance

Averages are useful, but they can hide the failure mode that matters most. If a system is correct 95 percent of the time on routine cases but unreliable on exceptions, its value depends on whether the workflow can detect and safely route the remaining 5 percent.

The relevant question is not merely “How accurate is the model?” It is:

Where does the model fail, how visible are those failures, and what happens when it does?

The coastline in fog

The frontier metaphor is useful because it changes how we think about rollout.

A wall would be easy to manage: everything on one side works; everything on the other does not. A coastline in fog is harder. Nearby tasks can lie on different sides of the line. Users cannot always see where the boundary is, particularly when the output remains confident and coherent.

This is why teams can have contradictory experiences with the same tool:

  • A customer-service team sees immediate gains
  • A legal team finds unreliable citations
  • Analysts accelerate routine reporting
  • Experts lose time validating complex edge cases

All may be correct. The system is operating across different parts of the frontier.

Map workflows, not job titles

The wrong unit of analysis is the job title. The right unit is the task and its verification environment.

Consider “financial analyst” as a role. It contains many different activities:

  • Collecting data from reports
  • Reconciling entities across systems
  • Creating routine commentary
  • Investigating unexpected variance
  • Deciding whether an anomaly is material
  • Defending a recommendation to senior stakeholders

Some are highly structured and easy to check. Others depend on tacit context, accountability, and judgment. AI will not affect every part of the role at the same speed.

A practical mapping exercise

For each workflow, score five dimensions:

  1. Source quality: Are the inputs accessible, structured, and governed?
  2. Task repeatability: Does the work follow a recurring pattern?
  3. Verification cost: Can the output be checked quickly and reliably?
  4. Error consequence: What happens if the output is wrong?
  5. Escalation path: Can uncertain cases be routed to a qualified owner?

Tasks with high-quality data, repeatable structure, cheap verification, manageable consequences, and clear escalation are strong candidates for automation.

Tasks with ambiguous data, costly verification, severe consequences, and unclear ownership need a different design: AI-assisted investigation rather than unattended automation.

Build for graceful failure

The best enterprise AI systems do not assume the model will always be right. They are designed to fail visibly and safely.

That means:

  • Confidence alone does not decide action
  • Business rules flag known exceptions
  • Source evidence is available beside every material claim
  • Low-confidence or high-impact cases are escalated
  • Teams monitor failure patterns over time
  • Feedback improves the workflow, not only the prompt

The objective is not to eliminate every error before deployment. It is to make the remaining error legible, bounded, and recoverable.

The AIxtract approach

AIxtract is most useful when it helps organisations see the frontier rather than pretend it does not exist. By connecting outputs to underlying data and evidence, it makes it possible to deploy AI where it is reliable, identify exceptions where it is not, and continuously refine the boundary.

The question is never simply whether AI works. The question is whether the workflow knows what to do when it does not.


Adapted from the book Cheap Thinking: What AI Makes Abundant, What It Makes Scarce, and Who Captures the Difference by Daoyuan Li, PhD.