Intelligent Agent

Intelligent Agent

Definition: An intelligent agent is any entity that perceives its environment through sensors and acts upon that environment through actuators in pursuit of a goal or performance objective. It is the unifying abstraction of classical AI — the same framing covers a thermostat, a chess program, a warehouse robot, and a large language model that calls tools. What separates a “smart” agent from a bare control loop is not autonomy alone but the sophistication of the mapping from percepts to actions: how much the agent remembers, models, plans, and learns before it commits to an action. Russell and Norvig’s textbook formalization of this idea is the closest thing AI has to a single organizing theory, and it predates — and now underlies — the modern wave of “AI agents” built on language models.

How It Works

The Agent Function and Rational Action

  • Formally, an agent’s behavior is described by an agent function f:P∗→Af: P^* \to A, mapping every possible percept sequence (everything the agent has ever sensed, in order) to an action.
    • This is a mathematical spec, not code — in principle it could be written as a giant lookup table, though for any nontrivial environment that table would be astronomically large.
  • The agent program is the concrete piece of software or circuitry that runs on some physical or virtual architecture and implements (an approximation of) that function.
    • Architecture is the hardware or runtime the program executes on: a CPU and motors for a robot, a server and API layer for a chat agent.
    • The same agent function can, in principle, be realized by very different agent programs — a lookup table and a learned neural policy can implement the same input-output mapping while working nothing alike internally.
  • A rational agent is one that, for each possible percept sequence, selects the action expected to maximize its performance measure, given the evidence in that sequence plus whatever built-in or learned knowledge it has.
    • Rationality is judged at decision time, on expected value — not in hindsight, and not against some absolute notion of correctness.
  • Rationality is not omniscience. A rational agent can still make a “wrong” move if the environment surprises it; it is judged on the expected outcome given what it could plausibly have known.
    • A rational poker agent can fold a hand that would have won — that doesn’t make the fold irrational if the odds favored folding at the time.
  • Rationality is also not the same as human-likeness. A rational vacuum robot doesn’t need to reason like a person; it needs to reliably maximize its performance measure within its environment.
    • This is precisely why the Turing Test — a test of human-likeness — is a different and narrower question than agent rationality.
  • The vocabulary has old roots: Alan Turing’s 1950 paper first framed machine intelligence as behavior, John McCarthy coined “artificial intelligence” in 1956, and Herbert Simon’s work on bounded rationality gave the field its economic grounding.
    • Russell and Norvig’s Artificial Intelligence: A Modern Approach later packaged all of this into the agent formalism — agent function, PEAS, environment properties — that the field still teaches today.
  • Russell and Norvig further separate ideal rationality (always taking the objectively correct action, with unlimited compute) from calculative rationality (eventually computing the correct action, possibly too late to matter) and bounded rationality (doing the best that’s achievable within real time and resource limits).
    • Almost every deployed agent — from a thermostat to a production LLM agent — only ever achieves the third kind, which is why bounded rationality, not ideal rationality, is the practical design target.

Bounded Rationality and Satisficing

  • Herbert Simon’s key insight, decades before modern AI, was that real agents never have unlimited time, memory, or computation to find the optimal action.
    • Computing a perfect PEAS-optimal decision for a rich environment can be intractable, even when the environment’s rules are fully known.
  • Bounded rationality describes agents that reason well given their actual computational limits, rather than agents that reason perfectly given unlimited resources.
    • A chess engine that searches to a fixed depth and evaluates with a heuristic function is bounded-rational; true optimality would require searching the entire game tree.
  • Satisficing — choosing the first action that is “good enough” against a threshold, rather than searching exhaustively for the best possible action — is the practical strategy bounded-rational agents use.
    • A goal-based route planner that stops at the first route under 20 minutes, instead of exhaustively finding the single fastest route, is satisficing.
  • This distinction matters enormously for engineering: a design that assumes unbounded compute (search the whole space) will not survive contact with a real-time, resource-constrained deployment.
    • It is also why “just make the model bigger” or “just search deeper” is not a universal fix — bounded resources are a permanent constraint, not a temporary one that better hardware eventually erases.
  • Anytime algorithms formalize satisficing further: they can be interrupted at any point during deliberation and still return a valid, improving answer.
    • This is exactly what a bounded-rational agent needs under a hard real-time budget — a planner that’s cut off mid-search should still hand back its best answer so far, not nothing at all.

The Perceive-Think-Act Cycle

  • Every agent, no matter how simple, runs some version of a loop: sense the world, update or consult internal state, decide on an action, execute it, and repeat as the environment changes in response.
  • Purely reactive agents collapse “think” into a direct lookup from percept to action, with no memory and no deliberation step at all.
    • This is fast and robust in simple environments, but it cannot handle any situation where the correct action depends on something outside the current percept.
  • Deliberative agents insert real computation — search, planning, inference — between perception and action, at the cost of latency.
    • This tradeoff is exactly why a self-driving car’s obstacle-braking reflex and its route-planning module run as separate subsystems with very different time budgets.
  • Most production systems are hybrid: a fast reactive layer handles time-critical responses while a slower deliberative layer handles planning.
    • Think of a warehouse robot that reflexively stops for an unexpected obstacle while a separate planning layer recomputes its overall path.
  • The loop only works if actuation actually changes the environment in a way sensors can later detect — an agent with a broken feedback path is flying blind regardless of how smart its “think” step is.
    • Robotics calls this the sense-plan-act pipeline; the naming differs by subfield, but the structure — perceive, decide, act, observe the result — is identical to the general agent loop.

Environment Properties That Shape Agent Design

The right agent architecture is determined almost entirely by the environment it operates in, not by taste. The key dimensions:

PropertySpectrumWhat It MeansDesign Impact
ObservabilityFully vs. partially observableCan sensors see the complete relevant state at every step?Partial observability forces the agent to keep internal state, since percepts alone are ambiguous
DeterminismDeterministic vs. stochasticDoes an action from a given state always produce the same next state?Stochastic environments require reasoning about probabilities, not single outcomes
DynamicsStatic vs. dynamicCan the environment change while the agent is deliberating?Dynamic environments punish slow planners; they need bounded thinking time or reactive fallbacks
ContinuityDiscrete vs. continuousAre states, time, and actions countable or continuous-valued?Continuous environments usually need control theory or function approximation, not lookup tables
Episodic vs. sequentialEpisodic vs. sequentialDo actions in one episode affect future episodes?Sequential tasks require planning and credit assignment across time; episodic tasks do not
Agent countSingle vs. multi-agentAre there other agents whose actions affect the outcome?Multi-agent settings introduce competition, cooperation, and game-theoretic reasoning
KnowledgeKnown vs. unknownAre the environment’s rules and physics known in advance?Unknown environments require exploration and learning before exploitation is safe
TimingReal-time vs. offlineMust the agent act within a hard time budget?Real-time constraints cap how much search or inference the “think” step can afford to do

A chess-playing agent sits at fully observable, deterministic, static, discrete, sequential, multi-agent, and known — which is exactly why it was one of the earliest AI problems cracked. A self-driving car sits near the opposite end on almost every axis, which is exactly why it remains one of the hardest.

Two more quick classifications build the same intuition:

  • Poker-playing agent: partially observable (opponents’ hands are hidden), stochastic (card draws), sequential, discrete, multi-agent, known rules.
  • Email spam filter: partially observable (sender intent is never directly visible), stochastic, largely episodic (each message judged mostly independently), discrete, effectively single-agent, and only partly known — spammers keep changing tactics.
  • Self-driving taxi (from the PEAS example above): partially observable, stochastic, dynamic, continuous, sequential, multi-agent, and known-ish — traffic law is known, but other drivers’ intentions are not.

Why It Matters

  • It’s the field’s organizing abstraction. Search engines, recommendation systems, robots, and chatbots all reduce to the same percept-decide-act frame, letting researchers compare wildly different systems on common terms.
    • A search engine and a Mars rover look nothing alike, yet both are fully described by naming their PEAS and their agent tier.
  • It explains the current “AI agents” wave. When an LLM is described as “acting as an agent,” it means the model has been wired into a perceive-think-act loop — retrieved context as sensor, tool calls as actuators — turning a stateless model into something with classical agent structure.
    • Without that wiring, a language model is just a function from text to text; it has no percepts, no actuators, and no loop at all.
  • PEAS analysis forces scope discipline. Naming the performance measure, environment, actuators, and sensors before writing code catches ambiguous requirements early, the same way a spec catches bugs before implementation.
    • Teams that skip it tend to discover the missing piece only after deployment, when the agent behaves “correctly” by an unstated metric nobody actually wanted.
  • Environment properties predict which techniques will work. Knowing an environment is partially observable and stochastic tells you immediately that belief states and probabilistic reasoning are required, not a simple lookup table.
    • This turns architecture selection into a checklist exercise instead of a matter of engineering taste.
  • Rational-agent framing decouples “smart” from “human-like.” It lets researchers evaluate an agent purely on outcomes relative to its performance measure, sidestepping unproductive debates about whether it “really understands” anything.
    • This is also why a narrow chess engine can be more “rational,” in this technical sense, than a human grandmaster within its own environment.
  • It underwrites multi-agent and economic theory. Auction design, traffic systems, and game-theoretic AI all build directly on the single-agent formalism, extended to strategic interaction between multiple rational agents.
    • Mechanism design — building rules so that self-interested rational agents produce a good collective outcome — is this idea applied to markets and auctions.
  • It grounds AI safety and alignment work. Misalignment is most precisely described as an agent optimizing its actual performance measure or reward signal instead of the one its designers intended — a framing that only makes sense inside agent theory.
    • Specification gaming, reward hacking, and goal misgeneralization are all agent-theory terms before they are safety terms.
  • It structures robotics engineering. Real robots are built as an explicit sense-plan-act pipeline, or a reactive/hybrid variant of it, and debugging a robot usually means isolating which stage of that pipeline is failing.
    • “Is this a perception bug or a planning bug?” is one of the first diagnostic questions any robotics engineer asks.
  • It shapes product architecture decisions. Whether to build a real-time reactive system or a slower, more deliberative one is an environment-driven engineering tradeoff, not a stylistic preference.
    • A fraud-detection agent screening transactions in milliseconds cannot afford the same deliberation budget as an overnight batch-planning agent.
  • It’s directly testable and gradeable. Because a performance measure is explicit, agents can be benchmarked — win rate, cumulative reward, task success rate — in a way that vaguer notions of “intelligence” cannot be.
    • This is what makes leaderboard-style evaluation of agentic LLM systems possible at all: a defined task plus a defined success metric.

The PEAS Framework

PEAS — Performance measure, Environment, Actuators, Sensors — is the standard checklist for specifying an agent before building one. Skipping it is the single most common reason agent projects drift: without an explicit performance measure, “success” becomes whatever the last person in the room decided it meant.

ComponentQuestion It AnswersSelf-Driving Taxi Example
Performance measureHow is success scored?Passenger safety, obeying traffic law, minimizing trip time, ride comfort, fuel/energy efficiency, fare earned
EnvironmentWhat is the agent operating in?Roads, intersections, other vehicles, pedestrians, weather, traffic signals, the passenger
ActuatorsWhat can the agent do?Steering, accelerator, brakes, turn signals, horn, gear shift, display/speaker for passenger communication
SensorsWhat can the agent perceive?Cameras, lidar, radar, GPS, speedometer, wheel encoders, microphone, engine/battery telemetry

A second, much narrower example makes the contrast clear:

ComponentRobotic Vacuum Example
Performance measureCleanliness per unit of time and battery used — not raw distance traveled, which a naive design might wrongly optimize for
EnvironmentA room with furniture, stairs, rugs, and possibly pets
ActuatorsWheels, rotating brushes, suction motor
SensorsBump sensors, cliff sensors, sometimes a camera or lidar for mapping

Notice how much smaller and more constrained the vacuum’s profile is compared to the taxi’s — that gap alone explains why household robots shipped a decade before autonomous cars did.

Two failure modes recur when teams write PEAS specs in practice:

  • Naming actuators without naming sensors. An agent cannot act rationally on a world it cannot perceive, no matter how capable its actuators are.
    • A robot arm with excellent motors but no force feedback will happily crush whatever it’s supposed to grip gently.
  • Writing a performance measure that’s easy to compute but only loosely correlated with what actually matters. “Distance driven” instead of “trips safely completed” tunes the agent toward the wrong target entirely.
    • This is the same trap as choosing a proxy KPI in a business dashboard: easy to measure is not the same as important to maximize.

Agent Architectures: Five Levels of Sophistication

Agents are conventionally classified by how much machinery sits between percept and action. Each tier subsumes the one before it and trades implementation simplicity for the ability to handle harder environments.

Agent TypeDecision BasisKeeps Internal State?Handles Partial Observability?ExampleKey Limitation
Simple reflexFixed condition-action rules on the current percept onlyNoNoThermostat; a spam filter keying only off the current messageFails outright once the correct action depends on something not in the current percept
Model-based reflexCondition-action rules applied to an internal model of the unobserved parts of the worldYes (world model)YesA vacuum that remembers which rooms it already cleanedThe model can go stale if the environment changes faster than the agent updates it
Goal-basedSearch or planning toward an explicit goal stateYes, plus a goalYesGPS route planner; a puzzle-solving agentComputationally expensive; doesn’t distinguish between multiple ways of reaching the goal
Utility-basedMaximize an expected utility function over outcomes, not just reach a goalYes, plus a utility functionYesAn expected-utility trading algorithm weighing risk against returnRequires accurate utility values and outcome probabilities, which are often hard to estimate
Learning agentA performance element continuously improved by a critic and learning elementYes, and it evolves over timeYesA recommender retrained on click feedback; an RL game-playing agentNeeds safe, sufficient exploration; can learn the wrong lesson from biased or sparse feedback

These five tiers are Russell and Norvig’s canonical taxonomy, and the boundaries between them are conceptual rather than a strict engineering checklist — a real system can sit between two tiers, or apply a higher tier only to part of its decision-making. What matters is not memorizing the labels but recognizing which capability a given failure is missing.

Each tier earns its added complexity by fixing a specific failure of the tier below it:

  • Simple reflex agents fail the moment the world isn’t fully observable — model-based reflex agents fix this by remembering what they can’t currently see.
  • Model-based reflex agents still only ever chase whatever the rules encode — goal-based agents fix this by searching toward an explicit, changeable objective.
  • Goal-based agents treat every goal-satisfying path as equally good — utility-based agents fix this by scoring outcomes on a continuous scale, not a binary “reached it or not.”
  • Utility-based agents are only as good as the utility function and probabilities a human hand-designed — learning agents fix this by improving that function from experience instead of freezing it at design time.
  • None of the tiers are mutually exclusive labels for real systems: most production agents blend two or three tiers, using goal-based planning for the main task and a learning loop wrapped around it to tune parameters over time.

The utility-based tier is where real math enters agent design. Instead of asking “does this action reach the goal,” the agent asks which action maximizes expected utility:

EU(a∣e)=∑s′P(Result(a)=s′∣a,e)⋅U(s′)EU(a \mid e) = \sum_{s'} P(\text{Result}(a) = s' \mid a, e) \cdot U(s')

This sums, over every possible resulting state s′s', the probability of landing there times how good that state is. It’s what lets a utility-based agent choose the better of two goal-satisfying paths — faster but riskier versus slower but safer — instead of just the first one it finds.

A concrete worked comparison makes this precise. Suppose a routing agent must choose between two options that both eventually reach the destination:

  • Route A succeeds 95% of the time (utility 100 if successful, 0 if not): EU(A)=0.95(100)+0.05(0)=95EU(A) = 0.95(100) + 0.05(0) = 95.
  • Route B succeeds 99% of the time but is slower even on success (utility 90, 0 if not): EU(B)=0.99(90)+0.01(0)=89.1EU(B) = 0.99(90) + 0.01(0) = 89.1.
  • A utility-based agent picks Route A despite Route B’s higher raw success rate, because Route A’s expected value is higher — a distinction a plain goal-based agent, which only asks “does this reach the goal,” cannot make at all.

The learning agent tier adds a second loop on top of the perceive-think-act cycle, one that improves the decision-making machinery itself over time rather than just executing it:

The critic is what makes this different from unstructured trial and error — it measures performance against a fixed external standard rather than the agent’s own possibly-biased self-assessment. The problem generator is what stops the agent from only ever exploiting what it already knows; without it, a learning agent converges early and never discovers a strategy better than the first mediocre one it stumbled on.

Comparison

“Agent” gets used loosely in both AI research and product marketing. It’s worth separating it cleanly from adjacent terms.

ConceptPerceives Environment?Acts on Environment?Adapts / Learns?Key Distinction
Intelligent agentYes, via sensorsYes, via actuatorsOnly if it’s a learning agentThe full closed loop — the general category everything else here is a special case or a component of
Script / automation (RPA)Reads fixed inputs, no true sensing of a changing worldExecutes predefined stepsNoRuns a static sequence regardless of environment state; breaks the moment the world deviates from the script’s assumptions
Standalone AI/ML modelTakes a single input, no ongoing perception loopProduces an output, doesn’t itself act on the worldLearns during training, frozen afterward unless fine-tunedA model is a component an agent uses to decide; it isn’t itself wired into sensors and actuators
Rule engine / expert systemReads structured facts as inputOutputs a conclusion or recommendation, rarely acts directlyNo, unless paired with a separate learning layerReasons over a fixed knowledge base rather than a live sensor-actuator loop — see Expert Systems
Autonomous robotYes, physical sensorsYes, physical actuatorsDepends on architectureA robot is an agent whose environment happens to be the physical world — the theory is identical, only the hardware differs
Multi-agent systemEach member perceives its local environmentEach member acts, possibly affecting others’ environmentsDepends on member designExtends single-agent theory to strategic interaction — cooperation, negotiation, competition — among multiple agents

The most common confusion in practice: calling a single LLM API call with no loop, no persistent state, and no tool access an “agent.” By the definition used throughout this page, that’s just a model invocation — the agent label only applies once a genuine sense-decide-act loop exists around it.

This distinction is worth applying skeptically to vendor claims. Three quick questions separate a genuine agent from a relabeled script or model: Does it have real actuators that change something outside itself, not just a text output? Does it perceive the result of its own actions rather than running open-loop? And is there an explicit performance measure it’s actually being evaluated against, rather than a vague promise of “autonomy”? A product that fails all three is agent-branded, not agent-shaped.

Real-World Use Cases

  • Self-driving vehicles (highway autopilot systems, robotaxi fleets) — the canonical goal-/utility-based agent operating in a partially observable, dynamic, stochastic, multi-agent environment.
    • Waymo and similar systems run separate perception, prediction, and planning modules that together implement exactly this loop at highway speed.
    • Most production systems still keep a human fallback specifically because full autonomy across every corner case of this environment remains unsolved.
  • Robotic vacuum and lawn-mowing robots — model-based reflex agents that maintain an internal map to avoid re-covering cleaned areas.
    • Later-generation models added lidar-based SLAM mapping, upgrading them from pure reflex to genuinely model-based agents.
    • Multi-floor map persistence across separate cleaning sessions is what distinguishes a real model-based design from a reflex agent with a memory bolted on.
  • Game-playing systems (chess and Go engines, real-time strategy game bots) — goal- and utility-based agents in fully or near-fully observable, adversarial, multi-agent environments.
    • AlphaGo layered a learning agent on top of a utility-based search, using a learned evaluation function instead of a hand-coded one.
    • Poker-playing agents add a further wrinkle: hidden opponent hands force genuinely probabilistic reasoning, not just adversarial search.
  • Algorithmic trading systems — utility-based agents balancing expected return against risk under genuinely stochastic, partially observable market conditions.
    • Execution algorithms specifically optimize a performance measure like implementation shortfall, not raw profit alone.
    • High-frequency strategies push the “think” step’s time budget down to microseconds, forcing nearly all reasoning to happen offline, ahead of time.
  • Smart thermostats and building HVAC controllers — simple-to-model-based reflex agents optimizing energy use against a comfort performance measure.
    • Learning thermostats add a critic that adjusts setpoints from observed occupant behavior over weeks.
    • Multi-zone commercial HVAC effectively runs as a small multi-agent system, one agent per zone, coordinating on a shared energy budget.
  • Warehouse and fulfillment robots (automated storage/retrieval systems) — goal-based agents coordinating path planning in a shared, dynamic, multi-agent floor space.
    • Fleets of these robots must also solve a multi-agent coordination problem to avoid gridlocking each other’s paths.
    • Amazon’s Kiva-derived robots largely solve this by never letting two robots plan through the same grid cell at the same moment.
  • LLM-based coding and research agents — a language model wrapped in a perceive-think-act loop where the sensor is retrieved context and the actuators are tool calls such as running code, searching, or editing files.
    • The “think” step here is the model’s forward pass over accumulated context, repeated every loop iteration.
    • Because the environment (a codebase) usually stays put between actions, these agents lean more goal-based than reactive, unlike a physical robot.
  • Customer-support and sales chat agents — goal-based agents that perceive a conversation, decide among scripted or tool-based actions, and act through APIs.
    • A well-designed one calls a real lookup-order tool rather than hallucinating an answer from the conversation alone.
    • Confidence-threshold escalation to a human functions as the agent’s safety net for exactly the cases its performance measure can’t confidently resolve alone.
  • Autonomous drones (inspection, delivery, agriculture) — model-based or utility-based agents navigating physically dynamic, partially observable airspace.
    • GPS-denied environments force these agents to fall back on visual-inertial odometry as a substitute sensor.
    • Airspace regulation acts as an additional, non-negotiable actuator limit: some actions are physically possible but legally forbidden.
  • Recommendation and ad-ranking systems — learning agents whose critic is user engagement feedback, continuously reshaping the performance element’s ranking behavior.
    • This is also where reward-hacking failures are most visible, as the system learns to maximize engagement rather than genuine user value.
    • Production A/B testing is, functionally, the problem generator from the learning-agent diagram — deliberately trying non-greedy options to keep learning.

Common Pitfalls

  • Conflating autonomy with intelligence. A thermostat is, technically, an agent — it perceives and acts without human intervention — but almost no one would call it intelligent. Autonomy is necessary, not sufficient.
    • Marketing an agent as “intelligent” purely because it runs unattended sets expectations the underlying decision logic can’t meet.
  • Skipping the PEAS analysis before building. Teams that jump straight to an architecture without naming the performance measure, environment, actuators, and sensors routinely discover mid-project that they built the wrong loop for the wrong environment.
    • The fix costs an afternoon of specification; the bug it prevents can cost months of rework.
  • Optimizing the wrong performance measure. An agent tuned against a proxy metric — clicks, time-on-page, distance traveled — will cheerfully maximize the proxy at the expense of what actually mattered, a specific case of the reward-hacking problem central to AI Alignment.
    • A support agent scored on “messages sent” will learn to chatter rather than resolve.
  • Assuming the environment is static when it isn’t. A model-based agent whose internal world model is never refreshed will confidently act on stale information the instant the real environment shifts underneath it.
    • A pricing agent trained on last quarter’s demand curve will misprice the moment market conditions move.
  • Ignoring partial observability. Designing as if the agent can see the full state, when its sensors genuinely can’t, produces brittle behavior that looks fine in a demo and fails the first time an unobserved variable matters.
    • Demos are usually run in clean, fully-observable conditions precisely because that’s when the flaw stays hidden.
  • Underestimating multi-agent interaction effects. Two individually rational agents — two trading algorithms, two delivery-routing bots — can jointly produce irrational system-level outcomes, like flash crashes or gridlock, that no single-agent analysis would predict.
    • Testing each agent in isolation, without the other agents present, will not surface this class of failure at all.
  • Treating a raw LLM call as an agent. A model that only produces text isn’t wired into a real perceive-act loop; it becomes an agent only once given actual sensors and actuators with a genuine feedback path back to it.
    • Calling a single unlooped prompt-response exchange an “agent” is a category error that inflates what the system can actually do.
  • Mishandling the exploration/exploitation tradeoff in learning agents. Too little exploration and the agent gets stuck on a mediocre early strategy; too much and it never reliably exploits what it has already learned.
    • Production systems usually cap exploration with a small, bounded budget — a few percent of traffic — rather than leaving the ratio unconstrained.
  • No fallback for actuator or sensor failure. Designs that assume perfect execution of every action break catastrophically the first time a real actuator jams or a sensor returns garbage instead of raising an error.
    • Production agents need an explicit “I’m uncertain, escalate” action, not just a set of task-completing ones.
  • Choosing an architecture before understanding the environment’s properties. Deploying a learning agent into a fully known, static, deterministic environment — or a simple reflex agent into a stochastic, partially observable one — is a mismatch that no amount of tuning fixes; the environment should dictate the architecture, not the reverse.
    • The fix is always the same order of operations: classify the environment first, using the property table above, then pick the cheapest agent tier that can actually handle it.

Most of these ten failures trace back to one of two root causes: an unstated or wrong performance measure, or a mismatch between the agent’s architecture and the environment’s actual properties. Catching either one at design time, with an explicit PEAS specification and an honest environment classification, is cheaper than catching it in production.

Example

Consider a customer-support system built to handle refund requests over chat, wired up as a genuine intelligent agent rather than a scripted bot. Its performance measure is a weighted combination of resolution speed, policy compliance, and customer satisfaction score. Its environment is the live chat conversation plus the company’s order and refund-policy databases. Its actuators are a set of callable tools — issue-refund, apply-discount, escalate-to-human, send-message — and its sensors are the incoming customer messages plus whatever the API calls to those tools return. This is a full PEAS specification, and skipping any one piece of it, such as never defining what “success” means numerically, is exactly the kind of gap that leads teams to ship an agent nobody can evaluate.

On a percept-think-act cycle, the agent reads an incoming message — “my order never arrived, I want a refund” — and updates its internal state with the extracted order ID and complaint category before deciding on an action. Refund decisions carry real financial and policy consequences, so this sits closer to goal-based blended with utility-based reasoning than to a simple reflex lookup: the input space is far too varied for fixed rules to cover safely, and the agent needs to weigh the cost of a wrongful refund against the customer-satisfaction cost of a wrongful denial. It calls a lookup-order tool, which functions as an extension of its sensors into the company’s database, reasons over the returned shipping status, and then either calls issue-refund directly or escalate-to-human when the case falls outside its confidence threshold — the actuator step that actually changes the world.

Over time, a critic measuring real outcomes — whether escalated cases genuinely needed a human, whether issued refunds later turned out to be fraudulent — feeds signal into a learning element that adjusts the agent’s confidence thresholds and tool-selection behavior. This is precisely the learning-agent architecture described above, and it is also, functionally, what RLHF (Reinforcement Learning from Human Feedback)-style fine-tuning does to the underlying model across many such episodes. Strip away the chat interface and the language model, and what remains is the same agent theory that has described thermostats and mobile robots for decades — only the sensors, actuators, and scale have changed.

Mapped back onto the framework above, the whole scenario reduces to:

  • Performance measure: resolution speed, policy compliance, and satisfaction score, combined.
  • Environment: the chat conversation plus the order and refund-policy databases — partially observable, since fraud intent is never directly visible.
  • Actuators: issue-refund, apply-discount, escalate-to-human, send-message.
  • Sensors: incoming customer messages and tool-call responses from the order database.
  • Agent tier: goal-/utility-based at inference time, wrapped in a learning agent’s critic-and-feedback loop across deployments.

The broader lesson generalizes past this one example: any system billed as an “AI agent” can be audited by asking the same five questions — what is it actually optimizing, what can it perceive, what can it change in the world, what can’t it perceive, and does it get better over time or stay frozen at whatever behavior it shipped with. A system that can’t answer all five isn’t yet a fully specified agent, regardless of how capable its underlying model is.

Dig deeper