Skip to main content

Sentiment Analysis in NLP: Complete Guide with Python Code

NLP Sentiment Analysis: A Practical Guide from Lexicons to LLMs Oct 2, 2026 · @Syed Wahab Uddin Introduction: What Sentiment Analysis Is and Why It Matters Sentiment analysis is the NLP task of identifying the opinion, attitude or emotion expressed in text. At its simplest, it answers one question: is this text positive, negative or neutral? Also called opinion mining, it turns huge volumes of unstructured reviews, posts and messages into numbers a team can act on. Consider three everyday examples: "Delivery was quick and the packaging was perfect." is positive. "The app crashes every time I open my cart." is negative. "The order arrived on Tuesday." is neutral. A person labels these in a second. Doing it reliably for 50,000 reviews a day, in several languages, full of slang and sarcasm, is where NLP comes in. Why organizations invest in it Most of what customers think about a product is written down somewhere: app store reviews, support tickets, survey c...

How LLMs Generate Text: Tokens, Attention & Transformers Explained

 


How Large Language Models Generate Text: Tokens, Attention, and Transformers

Introduction

When you type a question into a chatbot and watch a coherent, relevant answer appear word by word, it can feel almost magical. Underneath that experience, though, is a specific, well-understood (if intricate) piece of machinery: a large language model repeatedly predicting, one small piece at a time, what text is most likely to come next, given everything that has come before. There is no separate "understanding module" and "writing module" — the entire behavior of the model, from answering factual questions to writing poetry to debugging code, emerges from this single, repeated act of next-piece prediction, executed by a neural network with a very particular architecture called a transformer.

This article walks through that machinery from the ground up. We'll start with tokens, the small units of text a model actually operates on, and see why text has to be broken up this way before a model can process it at all. We'll then look at how those tokens get converted into the numerical vectors a neural network can work with, and how the transformer architecture's key innovation — a mechanism called attention — lets the model figure out which parts of a passage are relevant to which other parts. From there we'll look at how a trained model actually generates text one token at a time, what "training" actually means for a model like this, and finally why, despite all of this being "just" next-token prediction, the resulting systems can carry out sophisticated reasoning, hold coherent conversations, and generalize to tasks they were never explicitly trained on.

1. Why Text Has to Be Broken Into Tokens

Neural networks, at their core, are mathematical functions that operate on vectors of numbers. Before any of the more sophisticated architecture we'll discuss can do anything at all, raw text — a sequence of characters — has to be converted into some numerical form the network can actually process. The specific way this conversion happens is called tokenization, and the resulting numerical units are called tokens.

Why not just use individual characters?

The most naive approach would be to treat each individual character as a token — feed the model one number per letter. This has an appealing simplicity, but it runs into a serious practical problem: it makes every sequence extremely long. A single English word averages around four or five characters, meaning a paragraph of a hundred words would require several hundred individual character-tokens to represent. Since the amount of computation a transformer needs to process a sequence tends to grow faster than proportionally with sequence length (as we'll see when we discuss attention), processing text at the character level is computationally expensive and makes it harder for the model to relate distant parts of a long document to each other, since they end up separated by an enormous number of intervening tokens.

Why not just use whole words?

The opposite extreme — treating each whole word as a single token — solves the length problem but introduces a different one: vocabulary size and the problem of unseen words. English alone has hundreds of thousands of distinct word forms once you account for different tenses, plurals, and so on, and a model would need a distinct numerical representation reserved for every single one of them, plus every word in every other language it needs to handle, plus every technical term, every product name, every possible typo. Worse, no matter how large this vocabulary is made, new words — invented product names, internet slang, technical jargon that didn't exist when the vocabulary was built — will inevitably appear that the model has genuinely never seen before, and a pure whole-word tokenizer has no good way to represent them at all.

The middle ground: subword tokenization

Modern language models use an approach that splits the difference, called subword tokenization. Rather than working with individual characters or whole words, the text is broken into pieces that are often smaller than a word but larger than a single character — common whole words might be a single token ("the," "is," "computer"), while less common or more complex words get split into a small number of meaningful pieces ("tokenization" might become "token" plus "ization," for instance). This is typically done using an algorithm (a common one is called byte-pair encoding) that is run once, ahead of time, over a huge amount of representative text, and works by starting with individual characters and repeatedly merging the most frequently co-occurring pairs of pieces into single, larger tokens, continuing until a target vocabulary size (often somewhere in the tens of thousands of distinct tokens) is reached.

This scheme gets the best of both worlds. Common words are represented efficiently as single tokens, keeping sequences reasonably short and computation manageable. And because the tokenizer works its way down to individual characters (or even individual bytes) as a fallback, any string of text whatsoever — including words the tokenizer has genuinely never seen before, made-up words, typos, or text in scripts the tokenizer wasn't specifically optimized for — can still always be represented as some sequence of tokens, even if that sequence ends up being less efficient (more, smaller tokens) for unusual or unfamiliar text than it would be for common, everyday language.

What this means in practice

A useful mental model, commonly cited as a rough rule of thumb for English text, is that a token corresponds to roughly three-quarters of a word on average — so a hundred-word passage of ordinary English prose translates to somewhere in the neighborhood of 130 tokens, though this varies depending on the specific tokenizer, the language being used (many tokenizers, having been built primarily on English-heavy training data, represent other languages considerably less efficiently, needing more tokens to cover the same amount of actual content), and how unusual or technical the vocabulary is. This is why, practically speaking, model context windows and pricing are usually described in terms of tokens rather than words or characters — tokens are the actual unit of currency the underlying model deals in, and understanding roughly how text maps to token counts is useful for anyone working closely with these systems, whether for managing costs, staying within a context window, or estimating how much of a long document can realistically be included in a single request.

2. From Tokens to Vectors: Embeddings as the Entry Point

Once text has been split into a sequence of tokens, each token still needs to be converted into a numerical vector before the neural network can do anything with it — a token identifier, which is really just an index into a fixed vocabulary list, is not by itself a meaningful numerical representation any more than an arbitrary word itself would be. This conversion happens through an embedding step, which is discussed in far more depth in the companion article on AI embeddings, but is worth touching on briefly here because it is the literal first step of how a transformer processes any input.

Each distinct token in the model's vocabulary has an associated vector, learned during training, stored in a large lookup table sometimes called the embedding matrix. When a piece of text is tokenized, each resulting token is looked up in this table, retrieving its corresponding vector, and the original sequence of tokens becomes a sequence of vectors — one per token — which is what the rest of the transformer architecture actually operates on.

There is one more piece needed before this sequence of vectors is ready for the transformer's main processing: some representation of position. A transformer, unlike some earlier neural network architectures for handling sequences, does not inherently process tokens strictly in order, one after another — its core mechanism, attention, treats the input more like an unordered set of vectors that all interact with each other simultaneously, rather than a strictly ordered chain. Since word order obviously matters enormously for meaning ("the dog bit the man" and "the man bit the dog" contain the exact same words but mean very different things), the model needs some explicit signal about each token's position in the sequence. This is handled by adding a positional encoding — another vector, specifically designed to encode a token's position — to each token's embedding vector, so the combined vector fed into the rest of the network carries information about both what the token is and where it sits within the sequence.

3. The Core Innovation: Self-Attention

With a sequence of position-aware token vectors in hand, we arrive at the piece of machinery that gives the transformer architecture its name and its power: the self-attention mechanism. Understanding attention well is the single most important step in understanding how modern language models actually work, so it's worth taking real time with it.

The problem attention solves

Language is full of long-range dependencies where the meaning of one word depends heavily on other words that might be far away in the sentence or paragraph. Consider the sentence: "The trophy didn't fit into the suitcase because it was too big." What does "it" refer to — the trophy, or the suitcase? A human reader resolves this instantly using world knowledge (trophies, not suitcases, are the kind of thing that would be "too big" to fit into something), but for a model to get this right, it needs some mechanism for the representation of the word "it" to be informed by, and connected to, the representations of both "trophy" and "suitcase" earlier in the sentence, even though several other words sit in between them.

Earlier neural network architectures for language, which processed text strictly one token at a time in sequence, struggled with exactly this kind of long-range dependency, because information about early words had to be carried forward step by step through many intermediate steps to reach a later word, and that information tended to degrade or get overwritten along the way, especially over long distances. Self-attention was designed specifically to solve this problem, by allowing every token's representation to directly and immediately incorporate information from every other token in the sequence, no matter how far apart they are, without needing to pass information through a long chain of intermediate steps.

How attention actually works

Self-attention works by having every token generate three derived vectors from its own representation, conventionally called a query, a key, and a value, each produced by multiplying the token's vector by a separate learned matrix of weights. Intuitively, and very roughly: the query vector represents what this token is "looking for" from other tokens in the sequence; the key vector represents what this token "has to offer" or advertises about itself to other tokens; and the value vector represents the actual content this token will contribute if another token decides its key is a good match for what it's looking for.

For a given token, its query vector is compared against the key vector of every other token in the sequence (including itself), typically via a mathematical operation called a dot product, which produces a raw score for how relevant each other token is to this one. These raw scores are then normalized (using a function called softmax) into a set of weights that sum to one — essentially, a distribution of "how much attention" this token should pay to each other token in the sequence. Finally, the token's updated representation is computed as a weighted sum of every token's value vector, weighted according to these attention weights.

Working through the trophy-and-suitcase example: when the model computes an updated representation for the word "it," the query generated by "it" would, if the model has learned to handle this kind of construction well, end up producing a high attention weight toward the key generated by "trophy" (or "suitcase," depending on context, or in genuinely ambiguous cases, some blend of both), meaning the value vector for "trophy" contributes heavily to "it"'s updated representation, effectively binding the pronoun to its likely referent. This is, in a real sense, exactly the kind of contextual disambiguation a human reader performs automatically, except here it emerges from a learned pattern of query-key-value matching rather than explicit, hand-coded grammatical rules.

Crucially, this entire process happens for every token in the sequence simultaneously and in parallel — every token computes its own query and compares it against every other token's key, all at once — rather than working through the sequence step by step. This is part of what makes transformers so much more efficient to train on modern hardware than the sequential architectures that preceded them: the heavy computation involved can be parallelized across many processors simultaneously, rather than being forced through a long, unavoidably sequential chain of steps.

Multiple heads: attending to different things at once

A single attention computation, of the kind just described, can only really capture one particular pattern of relatedness between tokens at a time. But language has many different, simultaneously relevant kinds of relationships — grammatical relationships (which word is the subject of which verb), reference relationships (which pronoun refers to which noun), topical relationships (which words relate to the same underlying theme), and many others, all operating at once within the same sentence. To capture this richness, transformers use what's called multi-head attention: rather than computing a single attention pattern, the model computes several — commonly a dozen or more — separate attention computations in parallel, each with its own independently learned query, key, and value matrices, allowing each "head" to specialize, through training, in picking up on a different kind of pattern or relationship. The results from all these separate heads are then combined back together into a single updated representation per token, which now reflects a rich blend of many different kinds of contextual relationships simultaneously.

Stacking layers

A single round of multi-head attention, followed by a modest amount of additional per-token processing (a small, straightforward neural network applied identically to each token's updated vector, sometimes called a feed-forward layer), constitutes one "layer" of a transformer. Real language models stack many of these layers on top of each other — commonly dozens, and in the largest models, well over a hundred — with the output of one layer becoming the input to the next. Each successive layer allows the model to build progressively more abstract and sophisticated representations: early layers tend to capture relatively local, surface-level patterns (basic grammatical relationships, common word pairings), while later layers, operating on the already-contextualized representations produced by earlier layers, are able to capture much more abstract and long-range patterns — the overall topic of a long passage, subtle implications, or relationships between ideas expressed many sentences apart. This layered, progressively-more-abstract structure is a big part of why deep transformer models are able to handle language with the sophistication they do, rather than being limited to shallow, purely local pattern matching.

4. From Representations to Predictions: How a Model Actually Generates Text

Everything described so far — tokenization, embedding, positional encoding, layers of multi-head self-attention — produces, for a given input sequence, a rich, contextualized vector representation for every token in that sequence. But a language model's actual job, at generation time, is to predict what token comes next. How does that final step work?

The final prediction step

After the input has passed through all of the transformer's layers, the model takes the final, fully-contextualized vector corresponding to the very last token in the current sequence, and passes it through one more learned transformation (often called the "language modeling head") that produces a score for every single token in the model's entire vocabulary — tens of thousands of numbers, one per possible next token, representing how likely the model thinks each one is to come next, given everything that came before. These raw scores are converted into an actual probability distribution (again typically using the softmax function mentioned earlier), so that every possible next token has an associated probability, and all of those probabilities sum to one.

Sampling: choosing the actual next token

Given this probability distribution over the entire vocabulary, the model needs to actually pick one specific token to output next. The simplest approach, called greedy decoding, is to always pick the single highest-probability token. This is deterministic and tends to produce fairly safe, predictable text, but it can also produce oddly repetitive or bland output, since it never takes any risks on a slightly-lower-probability but perhaps more interesting or varied continuation, and can sometimes get stuck in repetitive loops for exactly this reason.

In practice, most systems instead sample from the probability distribution, in a way that introduces some controlled randomness, governed by a setting often called "temperature." At a low temperature, the sampling process is pushed toward strongly favoring the highest-probability tokens, behaving similarly to greedy decoding. At a higher temperature, the probability distribution is "flattened" somewhat before sampling, giving lower-probability tokens a more meaningful chance of being chosen, which tends to produce more varied, creative, and less repetitive output, at some cost to strict coherence or factual reliability if pushed too far. Additional refinements — like restricting sampling to only the top handful of most likely tokens (called top-k sampling), or restricting it to the smallest set of tokens whose combined probability exceeds some threshold (called top-p or nucleus sampling) — are commonly layered on top of this basic temperature-based sampling to further control the balance between coherence and variety in the generated text.

One token at a time, autoregressively

Here is the crucial detail that ties this whole process together into actual, extended text generation: once the model has predicted and selected a single next token, that token is appended to the existing sequence, and the entire process repeats — the new, longer sequence (original input plus every token generated so far) is fed back through the model, producing a fresh probability distribution for the next token after that, which is again sampled, appended, and fed back in, over and over. This is what "autoregressive" generation means: each new token is generated conditioned on everything generated so far, one token at a time, in a loop, until either a special "end of sequence" token is generated (signaling the model considers its response complete) or some other stopping condition (like a maximum output length) is reached.

This has an important, somewhat counterintuitive implication: a language model does not plan out an entire response in advance and then write it down. It generates one token, looks at the (now slightly longer) sequence including that new token, generates the next one based on that, and so on — the appearance of a coherent, well-structured, multi-paragraph response emerges from many individual next-token predictions chained together, each one only "seeing" the tokens that came immediately before it, not some separately-formed overall plan sitting somewhere else in the model. This is part of why techniques like chain-of-thought prompting — encouraging a model to write out intermediate reasoning steps before producing a final answer — genuinely help with more complex tasks: by generating intermediate reasoning tokens first, the model effectively gives its own future self-generation steps more relevant context to condition on, since those earlier reasoning tokens become part of the sequence that subsequent tokens are generated in light of, functionally giving the model a way to "think through" a problem incrementally rather than needing to arrive at a complex answer in one single, unaided leap.

5. Training: Where Does the Model's Knowledge Actually Come From?

Everything discussed so far describes how a transformer processes and generates text once it has already been trained — but where do all the learned weights (the matrices used to compute queries, keys, and values in every attention head, the embedding table, the feed-forward layers, and everything else) actually come from?

Pretraining: learning to predict text at scale

The initial, and by far most computationally expensive, phase of training a large language model is called pretraining, and its core idea is remarkably simple to state, even though executing it at scale is an enormous engineering undertaking. The model is shown enormous quantities of text — a large fraction of the publicly available internet, books, code repositories, and other sources, amounting to hundreds of billions or even trillions of tokens — and, at every position in every piece of text, it is asked to predict what token comes next, exactly the same underlying prediction task described in the previous section, just now applied for the purpose of learning rather than generating.

Initially, before any training has happened, the model's weights are set randomly, and its predictions are essentially useless — a random guess across the entire vocabulary. But every time the model makes a prediction during training, that prediction is compared against what the token actually was in the real training text, and the difference between the two (formally, a "loss") is used, via an algorithm called backpropagation combined with gradient descent, to nudge every single one of the model's many weights very slightly in the direction that would have made that particular correct prediction slightly more likely. Repeated across trillions of individual predictions, across the entire enormous training corpus, this nudging process gradually shapes the model's weights into a configuration that has genuinely learned an enormous amount about the statistical structure of language — not just grammar and vocabulary, but facts about the world, patterns of reasoning, stylistic conventions, and much more, all absorbed purely as a side effect of getting progressively better at the single, simple task of predicting the next token in real text.

It's worth emphasizing just how much can be learned from this seemingly narrow objective. To predict the next token accurately across the enormous variety of text found across the internet and beyond, a model has to implicitly learn an awful lot: basic grammar and syntax, obviously, but also factual knowledge (to predict the next word in "the capital of France is ___" well, the model essentially has to have learned the fact), reasoning patterns (to complete a mathematical derivation or a logical argument correctly), stylistic conventions across many different genres and registers of writing, and even something resembling a model of how different characters or perspectives in a piece of text would plausibly think and speak. None of this is explicitly taught; all of it emerges as an implicit, necessary consequence of getting good at the underlying prediction task, applied at a scale of data and computation large enough for these more sophisticated capabilities to emerge.

Fine-tuning and alignment: shaping raw capability into a useful assistant

A model that has only gone through pretraining, while remarkably knowledgeable, is not automatically well-suited to being a helpful conversational assistant. Pretraining teaches a model to predict plausible continuations of text drawn from its training distribution, but "plausible continuation of internet text" and "helpful, honest, appropriately-formatted answer to a user's question" are not automatically the same thing — a purely pretrained model, given a question, might just as easily continue it with another related question, mimicking a pattern common in the kind of text (forum posts, FAQ lists) it was trained on, rather than actually answering.

This gap is addressed through additional rounds of training after pretraining, broadly called fine-tuning or post-training. One common technique, called supervised fine-tuning, continues training the model, but now specifically on a curated dataset of example conversations showing the kind of helpful, well-formatted responses that are actually desired, teaching the model, through further exposure to these specific examples, to reliably produce that style and structure of response going forward rather than whatever pattern happened to be most statistically common across its original, much broader pretraining data.

A further, often very impactful, stage involves training the model based on human (or, increasingly, AI-generated) feedback about which of several candidate responses to a given prompt is actually better — more helpful, more accurate, better formatted, safer — allowing the model to be further nudged toward producing the kinds of responses that are actually preferred, beyond what could be captured just from a fixed set of example conversations alone. This general family of techniques, often grouped under the term "alignment," is what turns a raw, purely next-token-predicting pretrained model into something that reliably behaves like a helpful, instruction-following assistant, and a substantial amount of the practical difference in quality and helpfulness between different language models, even ones built on broadly similar underlying architectures and trained on similar amounts of raw data, often comes down to differences in the quality and care put into this post-training and alignment process.

6. Why "Just" Predicting the Next Token Produces Such Sophisticated Behavior

A question that naturally arises from everything described above is: how can something as seemingly simple as repeatedly predicting the next token produce behavior as sophisticated as writing working code, solving novel logic puzzles, or holding a nuanced, multi-turn conversation? This is a genuinely interesting question, and while it remains an active area of ongoing research to fully understand why and how these capabilities emerge, a few observations help make it less mysterious.

First, "predicting the next token well" is a deceptively demanding task once you consider the full breadth of text a model like this is trained on. Predicting the next line of correct, working code requires something functionally very close to understanding how the code works. Predicting the next step in a mathematical proof requires something functionally very close to understanding the underlying mathematics. The task looks simple when stated abstractly, but achieving genuinely good performance at it, across the enormous variety of text found in a large-enough training corpus, requires the underlying model to develop internal representations and computational processes that go considerably beyond simple surface-level pattern matching or memorization — at a large enough scale of data and model capacity, getting good at the objective essentially forces the emergence of something that functions, for many practical purposes, like genuine understanding and reasoning ability, even though the training objective never explicitly demanded "understanding" as a separate, distinctly-specified goal.

Second, scale itself appears to matter in ways that are not always straightforwardly predictable in advance. Researchers have observed that certain capabilities — for instance, the ability to reliably perform multi-step arithmetic, or to follow complex, multi-part instructions — can appear fairly abruptly as models are scaled up in size and amount of training data, rather than improving smoothly and gradually the entire way. This phenomenon, often discussed under the term "emergent capabilities," suggests that some of what makes large models powerful is not just a linear accumulation of somewhat-better performance on the same simple task, but a qualitative shift in what the model becomes capable of doing at all, once enough scale and training have been applied — though the precise mechanisms behind why and when such jumps occur remain a genuinely active area of research and some debate.

Third, it's worth being appropriately careful not to overstate what is actually happening. A language model's "reasoning," produced through the mechanism described in this article, is fundamentally a learned statistical pattern of what kind of text plausibly follows other text, not a guaranteed, provably-correct symbolic reasoning process of the kind a traditional calculator or theorem-prover would carry out. This is exactly why language models can sometimes produce confident-sounding but factually incorrect statements (a phenomenon commonly called hallucination), or fail in surprising and sometimes strange ways on problems that seem, to a human, only slightly different from ones the model handles perfectly well elsewhere. Understanding the actual underlying mechanism — next-token prediction shaped by attention-based contextual processing, trained on enormous amounts of real text — helps explain both why these models are capable of so much, and why their capabilities, however impressive, remain fundamentally different in character from, and are not a substitute for, guaranteed, verifiable correctness in domains where that kind of certainty genuinely matters.

7. Context Windows and Their Practical Limits

One more practical concept worth understanding, tying directly back to the attention mechanism described earlier, is the idea of a context window: the maximum number of tokens a given model can take into account at once, spanning both the input it's been given and everything it has generated so far in its response. This limit exists for concrete architectural and computational reasons: the core self-attention computation, comparing every token against every other token in the sequence, grows in computational cost as the sequence gets longer, meaning there are real practical (and historically, quite severe) limits on how long a sequence a model can process at once, though substantial engineering and architectural improvements over time have made increasingly long context windows practical.

The context window has direct, tangible consequences for how these models are actually used. Anything not included within the current context window — a fact from a conversation that happened yesterday in a separate session, the contents of a huge codebase that can't fully fit into a single request, a lengthy reference document — is simply invisible to the model at generation time, no matter how relevant it might be, unless some external system (like the retrieval mechanisms discussed in the companion article on embeddings, or an agent's memory systems discussed in the companion article on AI agents) specifically fetches and reinserts the relevant portion of that information into the current context before generation happens. This is why techniques for managing what does and doesn't get included in a model's limited context — summarization, retrieval, careful curation of what's actually relevant to the task at hand — are such a central and recurring concern across essentially every serious application built on top of large language models, regardless of how large that context window happens to be for the particular model being used.

8. A Few Important Architectural Refinements

The description of self-attention and transformer layers above captures the core mechanism, but real, deployed models incorporate a number of refinements on top of this basic picture that are worth knowing about, since they show up constantly in discussions of specific models.

Causal masking: preventing the model from "seeing the future"

During training, a model is shown a large piece of real text all at once and asked to predict, at every single position simultaneously, what the next token is — an efficient way to extract many training signals from a single pass over a piece of text, rather than needing a separate pass for every position. But this creates an obvious risk: if the self-attention mechanism at a given position were allowed to attend to tokens that come later in the sequence, the model could simply "cheat" by looking ahead at the very token it's supposed to be predicting, learning nothing useful in the process. To prevent this, transformers used for text generation apply what's called a causal mask during attention: at any given position, a token is only allowed to attend to itself and to tokens that came before it in the sequence, never to tokens that come after. This masking is what makes the resulting model naturally suited to autoregressive, left-to-right generation, since it is trained, from the very beginning, under exactly the constraint it will actually operate under at generation time — never having access to anything beyond what has already been produced.

Layer normalization and residual connections

Training a network with dozens or even hundreds of stacked layers, of the kind found in modern language models, runs into practical difficulties if care isn't taken in how those layers are connected. Two techniques, generally used together throughout every layer of a transformer, help make this deep stacking actually trainable in practice. Residual connections allow each layer's output to be added to its own input, rather than fully replacing it, giving the network an easy "shortcut path" for information (and, critically, for training gradients flowing backward during learning) to pass through many layers without being forced through every single transformation along the way, which substantially eases the practical difficulty of training very deep networks. Layer normalization rescales the numerical values flowing through the network at various points to keep them within a well-behaved range, preventing values from growing uncontrollably large or shrinking to near-zero as they pass through dozens of successive layers, which would otherwise make training unstable or simply fail outright. Neither of these techniques is unique to transformers, but both are essential, largely invisible pieces of engineering that make training models as deep as modern ones practically feasible at all.

Efficient attention variants

The most literal, textbook version of self-attention requires comparing every token in a sequence against every other token, which means the computational cost grows roughly with the square of the sequence length — doubling the length of the input roughly quadruples the attention computation required. This becomes a serious practical bottleneck as context windows have grown from a few thousand tokens in early transformer-based models to hundreds of thousands or even beyond a million tokens in some current models. A substantial amount of applied research effort has gone into developing more computationally efficient variants of attention that approximate the same core idea — letting tokens dynamically attend to relevant other tokens — while avoiding the full quadratic cost, through techniques with names like sparse attention (only computing attention between certain structured subsets of token pairs rather than all of them), sliding window attention (restricting attention to a fixed-size local neighborhood around each token for most computation, occasionally supplemented with broader, less frequent global attention), and various clever, lower-level implementation optimizations that compute the mathematically exact same result as full attention but do so with dramatically better use of the way modern hardware actually moves and processes data. These innovations are a large part of why context windows have been able to grow as dramatically as they have in recent years without a proportionally enormous increase in the computational cost of using these models.

Mixture-of-experts architectures

A more recent architectural variation, used in some of the largest current models, is called mixture-of-experts. Instead of every single token passing through the exact same, single feed-forward layer at each position in the network (as in the more basic architecture described earlier), a mixture-of-experts layer contains many separate feed-forward "experts," and a small router component decides, for each individual token, which small subset of these experts (often just one or two, out of possibly dozens or more available) that particular token should actually be routed through. This allows the model's total parameter count — and, correspondingly, the sheer amount of learned knowledge and capability it can store — to grow enormously, while the actual computation required to process any given token stays much more limited than it would if every token had to pass through every single one of those parameters, since only a small, selected fraction of the total experts are actually invoked for any particular token. This has become an important technique for continuing to scale up model capability without a proportional, often prohibitive, increase in the computational cost of actually running the model.

9. Scaling Laws: Why Bigger (Usually) Means Better

One of the more consequential empirical discoveries in the development of large language models is that model performance tends to improve in a remarkably smooth and predictable way as three quantities are increased together: the number of parameters in the model (roughly, the number of individually-learned numerical weights, which corresponds to the model's raw capacity to store learned patterns), the amount of training data the model is exposed to, and the amount of computation spent on training. These relationships, often called scaling laws, have been characterized fairly precisely for a range of model sizes and training setups, and they have had an outsized influence on how the field has approached building more capable models: rather than primarily searching for fundamentally new and cleverer architectural tricks, a substantial amount of recent progress has come from an appropriately confident bet that simply training larger models on more data, guided by these empirically-derived scaling relationships, would reliably continue to produce meaningfully better results, and that bet has, by and large, proven correct across an impressively wide range of model scales.

An important, related, and somewhat subtler finding is that, for any given, fixed amount of available computation, there is an optimal balance to strike between how large a model to train and how much data to train it on — training a model that's too large on too little data, or conversely, training a model that's too small on a huge amount of data, both leave meaningful performance on the table compared to hitting something closer to the actual, computationally-optimal balance between the two. This has meaningfully shaped how model developers approach the practical decision of how to allocate a fixed, and often very large, training compute budget between model size and quantity of training data, rather than simply maximizing one or the other in isolation without regard for this trade-off.

It's worth noting that these scaling laws describe pretraining performance — specifically, how well a model gets at the underlying task of predicting the next token accurately — and while this generally correlates well with the kind of broader, more practically useful capabilities people actually care about (reasoning, following instructions, writing well, and so on), the relationship is not perfectly one-to-one, which is part of why the additional fine-tuning and alignment stages discussed earlier remain so important on top of raw pretraining scale alone, for actually turning a large amount of raw, scale-driven capability into a genuinely useful, well-behaved assistant.

10. Inference: Making Generation Fast Enough to Actually Use

Everything discussed so far describes what a model computes, conceptually, when generating text, but actually running that computation fast enough for a real, interactive product — where a user is waiting, in real time, to see a response appear — requires a substantial amount of additional engineering, generally referred to as inference optimization, distinct from the training process discussed earlier.

One especially important technique here is called key-value caching. Recall that autoregressive generation involves repeatedly feeding the entire sequence generated so far back through the model to predict just the next single token. A naive implementation of this would recompute the key and value vectors (from the self-attention mechanism described earlier) for every single token in the sequence, all over again, at every single generation step — an enormous amount of unnecessary, repeated computation, since the key and value vectors for tokens that were already processed in a previous step don't actually change on subsequent steps. Key-value caching avoids this waste by storing, and simply reusing, the previously computed key and value vectors for all earlier tokens, so that each new generation step only needs to compute the query, key, and value vectors for the single newly-added token, dramatically reducing the computation required per step compared to the naive alternative, and this technique is essentially universal across real, deployed generation systems.

Beyond caching, a range of further techniques — including reducing the numerical precision used to store and compute with the model's weights (a technique called quantization, trading a small amount of precision for substantial gains in speed and memory usage), batching multiple different users' requests together to make more efficient use of the specialized hardware these models run on, and specialized hardware and software specifically designed and optimized for exactly this kind of large-scale matrix computation — collectively make it practically feasible to serve responses from models with many billions of parameters at the speed and cost required for real, interactive consumer products, despite the sheer scale of computation conceptually involved in generating even a single response.

Summary

A large language model generates text through a specific, learnable process: raw text is broken into tokens, each token is converted into a numerical vector, and a stack of transformer layers — built around the self-attention mechanism, which lets every token's representation be directly informed by every other relevant token in the sequence, regardless of distance — progressively refines these vectors into rich, contextualized representations. From the final representation of the most recent token, the model produces a probability distribution over its entire vocabulary for what token should come next, one token is sampled from that distribution, and the whole process repeats, one token at a time, to produce extended, coherent text. All of this machinery is shaped, not by explicit hand-coded rules, but through training: an initial, massive pretraining phase where the model learns simply by trying to predict real text accurately across an enormous corpus, followed by additional fine-tuning and alignment stages that shape the resulting raw capability into a genuinely helpful, well-behaved assistant. The sophistication of the resulting behavior — reasoning through problems, writing functional code, holding coherent multi-turn conversations — emerges from this seemingly simple underlying mechanism because getting good at next-token prediction, across a large and varied enough body of real text, turns out to implicitly require something functioning very much like genuine understanding, even though understanding was never explicitly specified as the training objective. Grasping this underlying mechanism — tokens, embeddings, attention, and autoregressive generation, shaped through training — provides the foundation for understanding both what these systems are remarkably good at, and why they nonetheless remain fundamentally different from, and not a guaranteed substitute for, verifiable, symbolic correctness in the domains where that kind of certainty genuinely matters.

Comments

Popular posts from this blog

Multimodal AI Explained: How Text, Image & Voice Merge Into One Model

  The Rise of Multimodal AI: Text, Image, and Voice in One Model For most of the last decade, AI systems were narrow specialists: a language model that only understood text, an image classifier that only understood pictures, a speech recognition system that only understood audio. Getting these systems to work together meant stitching together separate pipelines, converting between formats at every handoff, and accepting the errors and awkwardness that came with each conversion step. That era is ending. Modern multimodal AI systems process text, images, audio, and increasingly video within a single unified model, reasoning across all of them together rather than treating each as a separate problem solved by a separate system. This guide explains what multimodal AI actually is, the real architectural difference between "bolted-together" and "natively unified" multimodal systems, where this technology is already changing real products, and an honest look at where it ...

PyTorch Explained: The Complete Guide to Deep Learning & Neural Networks in 2026

  PyTorch: The Complete Guide to Deep Learning's Most Popular Framework Introduction If you've trained a neural network, fine-tuned a language model, or experimented with a diffusion-based image generator in the last several years, there's a strong chance PyTorch was somewhere underneath it. Originally released by Facebook AI Research (now Meta AI) in 2016, PyTorch has grown from a research-focused alternative to established frameworks into the dominant tool in the deep learning world — powering everything from academic papers to some of the largest AI systems ever deployed in production. This guide takes a deep, practical look at PyTorch: what it is, why it was designed the way it was, how its core components fit together, and how to actually use it to build, train, and deploy real models. Whether you're completely new to deep learning or you've used other frameworks and want to understand what makes PyTorch different, this article will walk you through everythi...

AI Job Displacement 2026: What the Data Really Shows

  AI and Job Displacement: What's Actually Happening in 2026 Few questions about AI generate more anxiety, and more contradictory headlines, than what it's actually doing to jobs. One week brings a report of tens of thousands of layoffs attributed to AI; the next brings a forecast of net job creation once new AI-related roles are counted. Both can be true at once, describing different parts of a genuinely uneven, still-unfolding transition. This guide sets aside both the most alarmist and the most dismissive framings and works through what the actual 2026 data — from government labor statistics, corporate layoff tracking, and major research institutions — shows about where AI is displacing work, where it's mainly changing hiring rather than firing, and where the picture remains genuinely uncertain. Given how fast this data changes, treat the specific figures here as a snapshot of 2026, not a permanent verdict. 1. The Honest Headline: Displacement Is Real, Concentrated, ...