Large Language Model (LLM)
Large Language Model (LLM)
Definition: A Large Language Model is a neural network — almost always built on the Transformer Architecture — trained on massive corpora of text, and increasingly code and other modalities, to predict the next token in a sequence. At sufficient scale this simple predictive objective gives rise to broad language understanding, multi-step reasoning, translation, summarization, and code generation that were never explicitly programmed into the system. “Large” refers jointly to parameter count (typically billions to hundreds of billions of weights) and training corpus size (typically trillions of tokens). Modern LLMs are the computational engine behind most consumer-facing generative AI products, from chat assistants to coding copilots to autonomous agents.
How It Works
Pretraining: Next-Token Prediction at Scale
An LLM starts as a blank set of randomly initialized weights and is trained on a single, deceptively simple objective: given a sequence of tokens, predict the next one. Formally, the model learns a probability distribution over sequences by factoring it with the chain rule:
Training minimizes cross-entropy loss between the model’s predicted distribution and the actual next token, averaged across the corpus:
This is self-supervised — the “labels” are just the next word in the training text, so no manual annotation is needed at this stage. Pretraining corpora are assembled from several distinct sources, each contributing different kinds of knowledge:
- Web crawls — broad but noisy coverage of general knowledge, current events framing, and colloquial language.
- Books and long-form writing — narrative coherence, sustained argument structure, formal prose.
- Source code repositories — logical structure, syntax precision, and a large share of general reasoning ability, since code is unusually unambiguous training signal.
- Scientific papers and reference works — technical vocabulary and domain facts.
- Forums, Q&A sites, and dialogue transcripts — conversational patterns and informal reasoning traces.
Pretraining runs across thousands of GPUs or TPUs for weeks to months, and this single phase consumes the overwhelming majority of an LLM’s total compute budget — later stages are comparatively cheap. The output of pretraining is a base model: a raw completion engine that has absorbed grammar, facts, reasoning patterns, and stylistic conventions, but has no notion yet of being a helpful assistant — asked a question, it’s just as likely to continue it with more questions as to answer it.
Training Data Curation and Filtering
Raw web-scale text is not fed into pretraining as-is — it goes through substantial curation first, and the quality of that curation has a larger effect on final model quality than raw corpus size alone:
- Deduplication — near-identical documents and boilerplate are collapsed or removed, since repeated content otherwise wastes training compute and can cause the model to overfit to it.
- Quality filtering — heuristics and classifier models score documents for coherence and informativeness, downweighting or dropping low-value text like link farms and auto-generated spam.
- Benchmark decontamination — training data is checked against known evaluation sets so the model isn’t inadvertently trained on the exact questions it will later be tested on, which would invalidate those scores.
- Safety and legal filtering — content that’s illegal, clearly harmful, or raises licensing concerns is filtered out before training rather than being addressed only afterward at the alignment stage.
This curation step is one of the least visible parts of the pipeline from the outside, but practitioners widely regard it as one of the highest-leverage places to invest, since a cleaner, more diverse corpus can outperform a larger but noisier one at the same compute budget.
Tokenization and Vocabulary
Before any prediction happens, raw text is broken into tokens — the atomic units the model actually operates on. Most current LLMs use a subword scheme such as byte-pair encoding (BPE) or a close variant, which strikes a balance the two extremes miss:
- Character-level tokenization keeps the vocabulary tiny but makes sequences very long, which is expensive since compute scales with sequence length.
- Word-level tokenization keeps sequences short but produces a huge vocabulary and cannot represent words it has never seen.
- Subword tokenization merges frequently co-occurring character pairs into a fixed vocabulary of tens of thousands of entries, so common words collapse to a single token while rare words and typos fall back to smaller pieces.
This choice has consequences beyond raw model quality: because API pricing and context-window limits are measured in tokens rather than words or characters, the same sentence can cost meaningfully more in one language than another if that language’s script is underrepresented in the tokenizer’s training data. See Tokenization for the full mechanics.
The Transformer Backbone
Every mainstream LLM is a stack of transformer blocks, each combining self-attention with a position-wise feed-forward network. Self-attention lets every token look at every other token in the context and compute a weighted relevance score, which is what gives transformers their signature ability to model long-range dependencies:
Several structural choices make this deep stack trainable and expressive:
- Multi-head attention — several attention operations run in parallel subspaces, so the model tracks different relationship types (syntax, coreference, long-range topic dependency) at once instead of collapsing them into a single score.
- Positional encoding — attention is itself permutation-invariant, so position must be injected explicitly, via learned embeddings, sinusoidal functions, or rotary embeddings (RoPE) in most current models.
- Residual connections — each sub-layer’s output is added back to its input, which keeps gradients from vanishing across dozens to over a hundred stacked blocks.
- Layer normalization — rescales activations between sub-layers, stabilizing training at depth and scale.
See Attention Mechanism and Transformer Architecture for the full mechanics behind this backbone.
Model Architecture Variants
Not every LLM shares the same internal shape. Three design axes matter most in practice:
- Decoder-only vs. encoder-decoder — most modern generative LLMs are decoder-only: a single autoregressive stack handles both understanding the prompt and generating the response. Encoder-decoder designs, common in earlier machine-translation-style models, split the two roles into separate stacks connected by cross-attention.
- Dense vs. mixture-of-experts (MoE) — a dense model routes every token through every parameter on every forward pass. An MoE model instead holds many parallel “expert” sub-networks and routes each token to only a handful of them, so the total parameter count can be far larger than the compute cost per token would suggest — more capacity without a proportional inference cost:
- Multimodal extensions — many current LLMs extend the same transformer backbone to accept images, audio, or other modalities as input tokens alongside text, sharing the same attention and prediction machinery rather than bolting on a separate model. This overlaps with Computer Vision when the input modality is images: a separate encoder converts pixels into embedding vectors that get interleaved with text token embeddings, so the same attention layers reason jointly over “what’s in this picture” and “what the user is asking about it.”
Supervised Fine-Tuning and Alignment
A base model completes text; it does not converse. Supervised fine-tuning (SFT) retrains it on curated instruction-response pairs so it learns to behave like an assistant — following directions, adopting a consistent persona, and formatting answers usefully. On top of SFT, alignment techniques further shape behavior:
- RLHF (Reinforcement Learning from Human Feedback) — human or AI raters rank candidate responses, a reward model is trained on those rankings, and the policy is optimized against that reward signal (commonly with PPO) to increase helpfulness and honesty while suppressing unwanted outputs.
- Direct Preference Optimization (DPO) and related methods — newer approaches that optimize directly on preference pairs without training a separate reward model, simplifying the RLHF pipeline while targeting the same outcome.
- Constitutional / principle-based methods — the model critiques and revises its own outputs against a written set of principles, reducing reliance on large volumes of human-labeled preference data.
This overall process — SFT followed by preference-based alignment — is a specialization of the broader Fine-Tuning technique, applied specifically to reshape behavior rather than to inject new domain knowledge.
Parameter-Efficient Fine-Tuning
Retraining every weight in a multi-billion-parameter model for each downstream use case is expensive and, for most teams, unnecessary. A family of parameter-efficient methods gets most of the benefit at a fraction of the cost:
- LoRA (Low-Rank Adaptation) — freezes the original weights and injects small trainable low-rank matrices alongside them, cutting the number of trainable parameters by orders of magnitude while keeping quality close to full fine-tuning.
- Adapters — small trainable modules inserted between existing layers, leaving the pretrained weights untouched entirely.
- Prefix / prompt tuning — learns a small set of continuous “virtual tokens” prepended to every input instead of touching the model’s internal weights at all.
- Quantization-aware fine-tuning (e.g., QLoRA) — combines low-rank adaptation with reduced numerical precision, making it possible to fine-tune large models on a single consumer-grade GPU instead of a data-center cluster.
These techniques are also what make it practical for a single base model to be specialized into many task-specific or customer-specific variants without storing a full copy of the weights for each one.
Inference and Decoding
At deployment, the model generates autoregressively: a forward pass produces a probability distribution over the entire vocabulary for the next token, a decoding strategy picks one, the token is appended to the context, and the process repeats until an end-of-sequence token or length limit is reached. The decoding strategy shapes output character as much as the model weights do:
- Greedy decoding — always takes the highest-probability token; deterministic, but often repetitive or bland.
- Temperature sampling — flattens (higher temperature) or sharpens (lower temperature) the probability distribution before sampling, trading off creativity against reliability.
- Top-k / top-p (nucleus) sampling — restrict sampling to a plausible subset of the vocabulary, cutting off the long tail of unlikely tokens that temperature alone can still surface.
- Beam search — keeps several candidate sequences alive simultaneously to optimize joint sequence probability rather than greedily picking one token at a time; more common in translation-style tasks than open-ended chat.
- KV-caching — production systems cache previously computed key/value attention states so each new token doesn’t require recomputing attention over the entire prior context from scratch, which is what makes long-context chat responsive.
The Training-to-Deployment Pipeline
Why It Matters
- Powers most consumer-facing generative AI today: chat assistants, coding copilots, search summarization, and writing tools all sit on top of an LLM.
- Collapsed the cost of building NLP systems — one pretrained model plus prompting or light fine-tuning now replaces the separate bespoke pipelines that used to be built per task (translation, classification, summarization, entity extraction).
- Established the “foundation model” business model: labs compete primarily on pretraining scale and data quality, then monetize by licensing inference through APIs.
- Reframed programming itself — Prompt Engineering and Function Calling (Tool Use) let an LLM act as a reasoning layer that calls external tools, databases, and APIs, forming the backbone of modern Intelligent Agent and Multi-Agent System architectures.
- Reopened serious debate about Artificial General Intelligence (AGI) — capability jumps that came from scaling alone, without fundamentally new architectures, revived questions that had been mostly dormant since the symbolic-AI era.
- Created entire new disciplines around safety and governance: AI Alignment research, red-teaming, constitutional AI, model cards, and evaluation-as-infrastructure.
- Shifted compute economics industry-wide — training a frontier model now costs hundreds of millions of dollars and draws data-center-scale power, making compute and energy access a strategic resource, not just an engineering line item.
- Made grounding and retrieval first-class engineering concerns — Retrieval-Augmented Generation (RAG) exists specifically to patch an LLM’s static knowledge cutoff and reduce Hallucination.
- Commoditized faster than most expected — open-weight model families closed the gap with closed frontier labs within roughly a year of each major capability jump, compressing prices across the whole market.
- Changed what “knowing how to code” means for a huge swath of software work, as natural-language specification increasingly substitutes for hand-written implementation in first drafts.
Prompting and In-Context Learning
How a model is asked matters almost as much as which model is asked. A single set of frozen weights can behave very differently depending on how the prompt is structured, because the model conditions its entire output distribution on everything in the context window:
- Zero-shot prompting — asking the model to perform a task with instructions alone, no examples; relies entirely on knowledge and behavior baked in during pretraining and alignment.
- Few-shot prompting — including a handful of worked examples directly in the prompt; the model infers the pattern and applies it to a new case without any weight update, which is the clearest everyday demonstration of in-context learning.
- Chain-of-thought prompting — asking the model to reason step by step before giving a final answer, which measurably improves accuracy on multi-step arithmetic and logic tasks versus asking for the answer directly.
- System prompts — a persistent instruction layer set by the application rather than the end user, used to fix persona, constraints, and output format across an entire conversation.
- Structured-output prompting — constraining the model to emit JSON, XML, or another parseable format so its output can be consumed programmatically by downstream code, a prerequisite for most agentic and tool-calling pipelines.
This whole discipline is formalized as Prompt Engineering: treating the prompt itself as the primary lever for controlling model behavior, without touching the underlying weights at all.
Scaling Laws and Emergent Capabilities
Empirically, an LLM’s pretraining loss follows a remarkably smooth power-law relationship with three quantities: model size (parameters ), dataset size (tokens ), and compute (). A simplified form:
Loss keeps falling as any of these grow, but the relationship shows diminishing returns — each additional increment of capability requires an exponentially larger increment of compute. This is why frontier training runs have grown from single-digit millions of dollars to hundreds of millions in a few years for comparatively modest, if still meaningful, jumps in benchmark performance. A key refinement (the “Chinchilla” finding) showed many early large models were undertrained relative to their parameter count — for a fixed compute budget, there’s an optimal split between making the model bigger and feeding it more data, and many labs had been over-indexing on parameters alone rather than balancing the two.
Separately from the smooth loss curve, certain downstream capabilities appear to show up abruptly once a model crosses a scale threshold, rather than improving gradually alongside loss. This is the “emergent capabilities” phenomenon, and it’s genuinely debated whether these are true phase transitions in what the model can do, or an artifact of measuring with discontinuous metrics (like exact-match accuracy) that hide continuous underlying improvement until it crosses a scoring threshold. Commonly cited examples include:
- Reliable multi-step arithmetic and symbolic manipulation.
- Robust instruction-following across novel, previously unseen phrasings.
- In-context learning itself — performing a new task from a few prompt examples with no fine-tuning.
- Self-correction — noticing and fixing an error mid-generation when prompted to double-check its own work.
- Analogical transfer — applying a pattern learned in one domain (say, legal argument structure) to a superficially unrelated one (say, debugging a piece of code) without being told the two are connected.
- Translating between language pairs that appeared only rarely, or never directly paired, in the training data, by composing knowledge learned from each language separately.
Context Windows and Model Classes
The context window — the maximum number of tokens a model can attend to at once, spanning both the prompt and its own generated output — and the parameter count together define what a model class is realistically useful for.
| Model Class | Typical Parameter Range | Typical Context Window | Deployment | Typical Use Cases |
|---|---|---|---|---|
| Small / edge | ~1B–3B | Small to moderate | On-device, mobile, embedded | Autocomplete, offline assistants, latency-critical or privacy-sensitive tasks, simple classification |
| Mid-size | ~7B–70B | Moderate to large | Self-hosted server, single/multi-GPU | General-purpose assistants, internal tools, cost-sensitive production apps, fine-tuned specialists |
| Frontier | Very large (often undisclosed / mixture-of-experts) | Large to very large | Cloud API, hyperscaler infrastructure | Complex multi-step reasoning, long-document analysis, advanced coding agents, research assistance |
A larger context window doesn’t just mean “can read a longer document” — it changes what workflows are viable at all. A large window lets an entire codebase, contract, or research paper collection sit directly in the prompt instead of being chunked and retrieved piecemeal, but every additional token also adds latency and cost, since self-attention’s compute and memory scale poorly with sequence length unless the architecture specifically optimizes for it.
Long-Context Techniques
Extending context beyond what a model was originally trained on isn’t free — several complementary techniques make it practical:
- RoPE scaling / positional interpolation — adjusts the rotary positional encoding math so a model trained on shorter sequences can generalize to longer ones without retraining from scratch.
- Sliding-window and sparse attention — instead of every token attending to every other token, each token attends to a limited local window or a learned sparse subset, cutting the quadratic cost at the price of some long-range precision.
- State-space and hybrid architectures — alternatives or complements to pure attention that compress history into a fixed-size state, trading some of attention’s flexibility for near-linear scaling with sequence length.
- Retrieval as a context-extension strategy — rather than stretching the model’s own window, Retrieval-Augmented Generation (RAG) keeps most of a large corpus out of the prompt entirely and injects only the relevant slice per query.
Cost-Per-Token Economics
Most commercial APIs price input tokens (the prompt) and output tokens (the generation) differently, and output is typically several times more expensive per token than input. The reason is structural, not arbitrary: processing a prompt is highly parallelizable — the model can compute attention over the whole input in one pass — while generating a response is inherently sequential, one token at a time, each depending on the last. This asymmetry is why techniques like prompt caching and shorter, more targeted output formats have an outsized effect on production costs compared to trimming the input side alone.
Evaluation and Benchmarks
Because LLMs are general-purpose, no single number captures “how good” one is — practitioners layer several evaluation types:
- Perplexity — a direct measure of how well the model predicts held-out text, a lower-is-better proxy for raw language-modeling quality inherited straight from the pretraining objective.
- Academic benchmarks — standardized test suites covering broad knowledge (multi-subject exam-style questions), coding correctness (functional test-passing on generated programs), and reasoning, used to compare models on a common scale.
- Human preference evaluation — pairs of model responses shown to human raters who pick the better one; this is the same signal used to train reward models during alignment, and it remains the closest proxy to real user satisfaction.
- LLM-as-judge evaluation — using a separate, typically stronger model to score or rank responses at a scale human raters can’t match, trading some reliability for throughput.
- Red-teaming and safety evals — adversarial probing specifically to surface harmful, biased, or policy-violating outputs before deployment, feeding directly into AI Alignment and AI Bias and Fairness work.
- Cost-normalized evaluation — comparing quality per dollar or per token rather than quality alone, since a marginally better model that costs several times more per request is often the worse engineering choice for a given product.
- Regression evals — a fixed internal test set re-run against every prompt or model change, specifically to catch silent quality drops that generic public benchmarks won’t reflect for an application’s particular use case.
Benchmark scores are also notoriously easy to game unintentionally through data contamination — if benchmark questions leaked into the pretraining corpus, a high score measures memorization, not the capability the benchmark claims to test. This is precisely why internal regression evals tied to a real application’s actual traffic patterns tend to be more trustworthy signals than any public leaderboard position.
Limitations and Open Research Problems
Common Pitfalls below covers mistakes practitioners make when using LLMs. This section is different: these are limitations inherent to the technology as it currently exists, not usage errors.
- Quadratic attention cost — standard self-attention’s compute and memory grow with the square of sequence length, which is why extending context windows requires architectural workarounds (sparse attention, sliding windows, state-space hybrids) rather than simply feeding in more tokens.
- Catastrophic forgetting — fine-tuning a model on a narrow new task can quietly degrade its performance on unrelated tasks it previously handled well, since weight updates are not scoped to the new skill alone.
- No persistent memory across sessions — a deployed model has no built-in mechanism to remember a previous conversation once it ends; any continuity has to be engineered externally by re-supplying prior context.
- Shallow generalization on truly novel problems — performance degrades on problems that are structurally unlike anything in the training distribution, even when they are logically simple, because the model is fundamentally pattern-matching over learned representations rather than deriving answers from first principles.
- Significant energy and hardware footprint — both training and, at scale, inference consume substantial electricity and specialized chips, which is an active area of efficiency research (better architectures, quantization, distillation) as much as a hardware problem.
- Difficulty with precise counting and fine-grained structure — tasks like counting characters in a word or tracking exact positions can fail because of how tokenization abstracts away sub-token detail, independent of the model’s broader reasoning ability.
Open-Weight vs. Closed-Weight Models
Beyond parameter count and context window, how a model’s weights are released shapes who can realistically use it and how:
| Aspect | Closed-weight (API-only) | Open-weight |
|---|---|---|
| Access | Hosted inference only, via API or app | Downloadable weights, run anywhere |
| Customization | Limited to prompting and provider-hosted fine-tuning | Full fine-tuning, quantization, architecture modification |
| Data privacy | Requests leave the user’s infrastructure unless a private-cloud deal exists | Can run entirely on-premises or air-gapped |
| Transparency | Training data and methods usually undisclosed | Weights inspectable; training-data disclosure still varies by release |
| Typical trade-off | Frontier capability with no infrastructure burden | More control and lower marginal cost at high volume, at the cost of self-hosting overhead |
Neither approach is strictly better. A startup with no ML infrastructure team gets to a working product fastest via a closed-weight API, while a company with strict data-residency requirements or extreme request volume often finds self-hosting an open-weight model cheaper and safer over time. The two ecosystems also feed each other — techniques proven inside closed frontier labs, such as better alignment methods or longer context handling, tend to surface in open releases within a development cycle or two, which is part of why the capability gap between the two has narrowed rather than widened over time.
Deployment and Serving Patterns
Getting a trained model to answer a single prompt in a demo is a different engineering problem from serving it reliably to millions of users. Production deployment involves choices that don’t show up in a model’s benchmark scores at all:
- Raw completion vs. chat completion APIs — most providers expose a chat-formatted interface (a list of role-tagged messages) rather than the raw text-continuation interface the base model actually implements internally; the wrapper handles turn structure and system-prompt placement automatically.
- Streaming vs. batch responses — interactive applications stream tokens back as they’re generated so the user sees progressive output instead of waiting for the full response, while offline or bulk-processing workloads favor batching many requests together for throughput.
- Prompt caching — reusing the computed attention state for a repeated prefix (a long system prompt or document, called many times) avoids redundant computation and meaningfully cuts both cost and latency for that class of workload.
- Multi-turn agent loops — an agent repeatedly calls the model, executes any tool calls it requests, feeds the results back in, and loops until the task is done, rather than a single prompt-response round trip.
- Distillation to smaller models — a large model’s outputs are used as training data for a smaller, cheaper model, trading some capability for a substantial drop in inference cost at high volume.
- Guardrails and output filtering — a layer around the raw model that checks inputs and outputs against safety and policy rules before anything reaches the end user, since alignment during training reduces but does not eliminate the need for runtime checks.
The right combination of these depends entirely on the workload — a real-time chat product optimizes for streaming latency, while a nightly batch job classifying a year of support tickets optimizes for total throughput and cost per token instead.
Comparison
LLMs are frequently confused with, or lumped together with, other approaches to building intelligent systems. The differences matter for choosing the right tool:
| Aspect | Large Language Model | Expert System | Task-Specific Supervised Model |
|---|---|---|---|
| Knowledge source | Statistical patterns learned from massive, broad text corpora | Hand-coded rules and facts from domain experts | Labeled examples for one narrow task |
| Adaptability | General-purpose; handles novel tasks via prompting or fine-tuning | Rigid; new knowledge requires manually adding rules | Narrow; only performs the task it was trained for |
| Transparency | Low — reasoning is implicit in billions of weights | High — rules are explicit and traceable | Moderate — depends on model type, often still opaque |
| Typical failure mode | Confident Hallucination on facts or logic | Brittle gaps where no rule covers the case | Silent degradation on inputs outside the training distribution |
| Development cost | Very high upfront (pretraining), cheap to reuse afterward | High ongoing cost to maintain the rule base | Moderate; requires labeled data per task |
In practice, the strongest production systems rarely pick just one of these — an LLM handling ambiguous natural-language input often calls a task-specific classifier for a well-defined sub-step, or defers to explicit business rules for anything regulatory or safety-critical, rather than trusting the LLM’s judgment end to end.
Real-World Use Cases
- Conversational assistants for general Q&A, drafting, and research (Claude, ChatGPT, and comparable products).
- Coding copilots that autocomplete, explain, refactor, and increasingly write and run entire features from a natural-language spec.
- Customer support deflection — answering common tickets automatically and routing only edge cases to humans.
- Enterprise document intelligence — summarizing contracts, extracting clauses, and flagging risk in legal and compliance review.
- Retrieval-augmented enterprise search, where an LLM answers questions grounded in a company’s internal knowledge base via Retrieval-Augmented Generation (RAG).
- Structured data extraction — turning unstructured text (emails, PDFs, call transcripts) into clean, schema-conformant records.
- Synthetic data generation used to train or distill smaller, cheaper downstream models.
- Machine translation and localization at a fluency level that increasingly rivals dedicated translation systems.
- Content moderation and classification at scale, flagging policy-violating text with nuanced context understanding rule-based filters miss.
- Agentic workflows that chain multiple LLM calls with tool use — browsing, filling forms, querying databases, writing and executing code — to complete multi-step tasks with minimal supervision.
Common Pitfalls
- Trusting fluent output as correct — an LLM can produce a confident, well-formatted, entirely wrong answer; fluency is not a proxy for accuracy.
- Treating the model as a live database — its knowledge is frozen at a training cutoff; without retrieval or tool access it cannot know about anything after that date.
- Ignoring context-window limits — content beyond the window gets silently truncated or pushed out, which can drop critical instructions or data without any explicit error.
- Over-indexing on parameter count — a bigger model is not automatically better for a given task; data quality, fine-tuning, and prompt design often matter more than raw scale.
- Skipping prompt-injection defenses — text retrieved from documents, web pages, or tool outputs can contain instructions the model may follow as if the user wrote them, especially in agentic and RAG pipelines.
- Assuming determinism — sampling temperature makes generation stochastic by default; identical prompts can produce different outputs on different runs unless temperature is set to zero and other sources of nondeterminism are controlled.
- Shipping without an evaluation harness — judging quality by spot-checking a few outputs instead of running systematic benchmarks or regression evals hides failure modes until they hit production.
- Conflating fluency with reasoning — chain-of-thought text that reads as a logical derivation can still arrive at a wrong conclusion; the explanation is generated, not necessarily the true computation the model performed.
- Underestimating cost and latency at production scale — autoregressive, token-by-token generation means cost and latency both scale with output length, which compounds fast across millions of requests.
- Ignoring the alignment tax — safety tuning that’s too aggressive can make a model unhelpfully cautious, refusing or hedging on entirely benign requests.
Related Terms
- Transformer Architecture
- Attention Mechanism
- Tokenization
- Fine-Tuning
- RLHF (Reinforcement Learning from Human Feedback)
- Hallucination
- Retrieval-Augmented Generation (RAG)
- Prompt Engineering
Example
A mid-size SaaS company wants an AI support agent that can answer questions about their product using their own documentation, not generic web knowledge. They start from a frontier LLM accessed via API — the pretraining and alignment work (next-token prediction over a huge corpus, then SFT and RLHF) has already been done by the model provider, so the company never trains a model from scratch. Instead, they build a Retrieval-Augmented Generation (RAG) layer: support docs are chunked, converted to Embeddings, and stored in a vector database, so at query time the system retrieves the most relevant passages and inserts them into the prompt alongside the user’s question.
Early testing surfaces a classic failure: the model occasionally states a subscription-tier limit that isn’t in any retrieved document — a hallucination generated from generic training knowledge about how SaaS pricing tiers usually work, rather than pulled from this company’s actual docs. The team tightens the system prompt to instruct the model to answer only from retrieved context and say “I don’t have that information” otherwise, and adds a lightweight fine-tune on their support conversation history to lock in tone and format. They also wire up Function Calling (Tool Use) so the model can query the live billing system directly for account-specific questions rather than guessing at an answer.
Six weeks after launch, the agent resolves most first-contact tickets without human involvement, but the team keeps a human-in-the-loop review queue for anything touching refunds or account cancellations. It’s a deliberate acknowledgment that even a well-grounded LLM is a probabilistic text generator, not a certified source of truth, and the cost of a confident wrong answer in that category is too high to automate away entirely.
Referenced by
- Agent SDKs and Frameworks
- AI Alignment
- Artificial General Intelligence (AGI)
- Artificial Intelligence MOC
- Attention Mechanism
- Computer Vision
- Cosine Similarity
- Explainable AI (XAI)
- Fine-Tuning
- Function Calling (Tool Use)
- Hallucination
- Intelligent Agent
- Knowledge Representation
- Model Context Protocol (MCP)
- Multi-Agent System
- Natural Language Processing (NLP)
- Prompt Engineering
- Retrieval-Augmented Generation (RAG)
- RLHF (Reinforcement Learning from Human Feedback)
- Tokenization
- Transformer Architecture
- Turing Test