Skip to main content

Sentiment Analysis in NLP: Complete Guide with Python Code

NLP Sentiment Analysis: A Practical Guide from Lexicons to LLMs Oct 2, 2026 · @Syed Wahab Uddin Introduction: What Sentiment Analysis Is and Why It Matters Sentiment analysis is the NLP task of identifying the opinion, attitude or emotion expressed in text. At its simplest, it answers one question: is this text positive, negative or neutral? Also called opinion mining, it turns huge volumes of unstructured reviews, posts and messages into numbers a team can act on. Consider three everyday examples: "Delivery was quick and the packaging was perfect." is positive. "The app crashes every time I open my cart." is negative. "The order arrived on Tuesday." is neutral. A person labels these in a second. Doing it reliably for 50,000 reviews a day, in several languages, full of slang and sarcasm, is where NLP comes in. Why organizations invest in it Most of what customers think about a product is written down somewhere: app store reviews, support tickets, survey c...

NLP Text Preprocessing: A Complete Guide with Python Examples




NLP Text Processing: A Practical Guide from Raw Text to Model-Ready Data

Introduction: Why Text Processing Is Where NLP Really Starts

Every NLP system you used today began with text processing. Your phone's autocorrect, your inbox's spam filter, a shopping site's search bar and the chatbot that answered your question all start by turning messy human writing into something a machine can count, compare and learn from.

Raw text is hard for computers. It arrives with typos, HTML tags, emojis, mixed languages, inconsistent spelling and creative punctuation. A review like "OMG new phone is fire but battery dies by 3pm smh" is perfectly clear to a person. To a program, it is just a sequence of characters until we process it.

Text processing is the set of steps that bridges that gap. It covers cleaning, normalization, tokenization, linguistic annotation and representation. Done well, it makes models more accurate, faster to train and easier to debug. Done badly, it quietly throws away the very signals your model needed.

Practitioners like to say "garbage in, garbage out." In NLP the saying gets very specific. A sentiment model that strips the word "not" during stop-word removal will read "not good" as "good." A search engine that never lowercases will treat "Python" and "python" as different words. A classifier trained on tidy news articles will stumble on tweets. Many NLP failures in production trace back to preprocessing choices, not to the model itself.

The field has also changed a lot. A decade ago, a typical pipeline leaned on heavy hand-crafted preprocessing: stemming, stop-word lists and sparse TF-IDF vectors. Today, transformer models such as BERT and modern large language models (LLMs) ship with their own learned tokenizers, and aggressive cleaning can actually hurt them. Knowing which steps to apply, and when to skip them, is now a core skill.

This guide walks through the full journey from raw text to model-ready data. You will learn:

  • How to clean and normalize text without destroying meaning
  • How tokenization works, from simple splitting to subword algorithms like BPE
  • When stop-word removal, stemming and lemmatization help, and when they hurt
  • How part-of-speech tagging, named entity recognition and parsing add structure
  • How text becomes numbers through bag-of-words, TF-IDF and embeddings
  • What changes when you work with transformers and LLMs
  • How to handle multilingual and code-mixed text
  • How to build a complete, reusable pipeline in Python

Each section includes practical examples and code you can adapt. Whether you are a student building a first project or a developer adding NLP features to a web app, the goal is the same: to make deliberate choices about your text, rather than copying a preprocessing recipe and hoping it fits.

The Text Processing Pipeline at a Glance

A text processing pipeline is an ordered series of steps that turns raw text into input a model can learn from. Which steps you need depends mostly on one question: will a classical model or a transformer consume the output?

[embed: node/2ddfcc82-a94b]

Both paths start the same way, with collection and cleaning, and then they split. A classical model such as logistic regression on TF-IDF features relies on you to normalize, tokenize, filter and vectorize, because it cannot learn those things itself. A transformer brings its own trained tokenizer and learns features from context, so most classic steps are skipped, and some would actively hurt it.

Three ideas hold on both paths:

  • Every stage is a trade-off. Each step removes some variation. The goal is to remove noise and keep signal, and what counts as signal depends on your task.
  • Order matters. URLs must be replaced before punctuation is stripped, contractions expanded before apostrophes vanish, and part-of-speech tags assigned before lemmatization.
  • The same pipeline must run everywhere. Whatever you do to training text, you must do identically to text at prediction time. Mismatches here are a leading reason models shine in a notebook and fail in production.

Real projects are rarely a single straight pass. You will loop back often: an error analysis reveals a new kind of noise, a cleaning rule turns out too greedy, or a new data source arrives with its own quirks. The sections below take each stage in turn, starting with cleaning raw text.

Collecting and Cleaning Raw Text

Cleaning removes everything that is not part of the language you want to model. The right recipe depends on where your text comes from, so look at real samples before writing any code.

Know your sources

Web pages carry HTML tags, navigation menus, cookie banners and scripts. PDFs carry broken lines, words hyphenated across line ends, running headers and page numbers. Social media brings mentions, hashtags, URLs and emojis, while support tickets carry signatures, quoted earlier replies and auto-generated footers.

A good habit is to print 50 random samples and read them. You will spot problems no tutorial predicts, such as one template sentence repeated across thousands of documents, or a field full of the word "NULL."

Fix encoding first

Encoding errors produce mojibake: garbled text like "café" instead of "café." It happens when UTF-8 bytes are decoded as Latin-1 or Windows-1252. Always open files with an explicit encoding (encoding="utf-8"), and use a library such as ftfy to repair text that was already mangled upstream.

Remove markup properly

Do not use regular expressions to strip HTML from full web pages. Use a real parser such as BeautifulSoup and its get_text() method, or a boilerplate remover such as trafilatura that keeps the main article and drops menus and footers. Regex is fine for small, predictable fragments, like <br> tags inside a database field.

Normalize Unicode

The same visible character can be stored in different ways. "é" can be one code point (U+00E9) or two ("e" plus a combining accent), and the two forms compare as different strings. Apply Unicode normalization: NFC for safe storage, or NFKC when you also want to fold compatibility characters, such as full-width "A" into "A" and the ligature "fi" into "fi."

Also watch for invisible characters such as zero-width spaces, non-breaking spaces and byte-order marks. They split tokens in odd places and make identical-looking strings unequal. Be careful, though: the zero-width non-joiner (U+200C) is meaningful in Persian and Urdu script, so do not strip it blindly.

Replace, don't just delete

You rarely need the exact URL, email address or username, but you often want to know one was there. Replace them with placeholder tokens like <URL>, <EMAIL> and <USER>. This keeps the signal (a link is a strong spam feature), shrinks the vocabulary and removes personal data.

Deduplicate and filter

Scraped and logged data is full of exact and near-duplicate documents. Duplicates bias training, waste compute and inflate test scores when copies leak across train and test splits. Hashing removes exact duplicates; MinHash with locality-sensitive hashing catches near duplicates.

Then drop documents that are too short to carry meaning, are mostly symbols or numbers, or are in a language you do not support. Language identification tools such as fastText's lid.176 model or langdetect make that last check cheap.

A starter cleaning function

import re
import unicodedata
from bs4 import BeautifulSoup

URL_RE = re.compile(r"https?://\S+|www\.\S+")
EMAIL_RE = re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.-]+\b")
MENTION_RE = re.compile(r"@\w+")
INVISIBLE_RE = re.compile(r"[\u200b\ufeff]")  # keep U+200C for Urdu/Persian

def clean_text(raw: str) -> str:
    text = BeautifulSoup(raw, "html.parser").get_text(" ")
    text = unicodedata.normalize("NFKC", text)
    text = URL_RE.sub(" <URL> ", text)
    text = EMAIL_RE.sub(" <EMAIL> ", text)
    text = MENTION_RE.sub(" <USER> ", text)
    text = INVISIBLE_RE.sub("", text)
    return re.sub(r"\s+", " ", text).strip()

One rule ties all of this together: keep the raw text stored beside the cleaned version. Cleaning choices change as projects mature, and you never want to re-collect data because an early regex was too greedy.

Text Normalization: Making Equivalent Text Look the Same

Normalization maps different surface forms of the same meaning onto one form. The payoff is fewer, more reliable features. The risk is erasing differences that matter, so treat every normalization step as a decision, not a default.

Lowercasing

Lowercasing ("Apple" to "apple") merges words that differ only by case. That shrinks the vocabulary and helps small datasets. It also destroys information: "Apple" the company and "apple" the fruit become one token, "US" the country becomes "us" the pronoun, and the anger in "THIS IS UNACCEPTABLE" disappears.

A practical rule: lowercase for classic bag-of-words models and search indexes, but keep case for named entity recognition and for cased transformer models. When unsure, try both and compare on a validation set.

Contractions, slang and abbreviations

Expanding contractions ("don't" to "do not", "I'm" to "I am") helps when negation matters and your tokenizer would split "don't" awkwardly. Chat and social text also needs slang maps: "u" to "you", "gr8" to "great", "idk" to "I don't know." Build these maps from your own data, because slang is specific to platforms, regions and age groups.

Numbers, dates and units

Should "$1,299", "1299 dollars" and "1.3k" be treated as the same? For topic classification the exact value rarely matters, so replacing digits with a <NUM> token or bucketing them reduces noise. For question answering, price extraction or medical text, the numbers are the point, so keep them and standardize their format instead.

Punctuation and repeated characters

Stripping all punctuation is a common beginner step and often a mistake. Question marks signal questions, exclamation marks signal emotion, and hyphens and apostrophes hold words together. A safer approach removes only true noise and squashes elongations: "soooo goooood" becomes "soo good." Keeping two repeats preserves the emphasis while collapsing endless variants into one.

Emojis and emoticons

Emojis carry sentiment and intent, so deleting them throws away signal. Convert them to text with the emoji library's demojize function instead: a red heart becomes :red_heart: and a crying face becomes :loudly_crying_face:. Classic models can then treat them as ordinary words.

Spelling and accents

Spell checkers such as SymSpell can fix typos in user text, but they also "correct" product names, slang and people's names. Measure their effect before adopting them. Accent stripping ("café" to "cafe") helps English search, but it damages languages where accents change meaning, such as French "où" (where) and "ou" (or).

Order matters

Normalization steps interact with each other. Expand contractions before removing apostrophes. Replace URLs before stripping punctuation, or "https://site.com" turns into "httpssitecom." Write each step as a small function and run them in an explicit, tested order.

Tokenization: Splitting Text into Units a Model Can Use

Tokenization splits text into tokens, the basic units every later step works with. Pick the wrong unit and nothing downstream can fully recover, so it deserves more thought than a quick text.split().

Sentence tokenization

Summarization, translation and many annotation tools first split documents into sentences. Splitting on periods fails fast: "Dr. Khan arrived at 3 p.m. on Jan. 5." has five periods and one sentence. Use trained splitters such as NLTK's Punkt (sent_tokenize) or spaCy's sentence segmenter, which learn abbreviations from data.

Word tokenization

Whitespace splitting is the simplest method, and it breaks on punctuation: "great!" and "great" become different tokens. Rule-based tokenizers in spaCy and NLTK handle punctuation, contractions ("don't" becomes "do" and "n't"), URLs and hyphenated words with language-specific rules.

Word-level tokens have two deep problems. Vocabularies explode, because a language has hundreds of thousands of word forms and every typo adds another. And any word missing from training becomes a single <UNK> (unknown) token, so the model learns nothing about "cryptocurrency" if it never saw it.

Character tokenization

Splitting into characters removes unknown words entirely, since every word is built from a small alphabet. The cost is very long sequences and little meaning per token; the model must learn spelling before semantics. Character models still shine for spelling correction, transliteration and very noisy text.

Subword tokenization

Subword tokenization sits in between and powers almost every modern model. Frequent words stay whole, while rare words split into reusable pieces. The vocabulary stays fixed, commonly tens of thousands of entries, and no word is ever fully unknown.

Three algorithms dominate:

  • Byte Pair Encoding (BPE) starts from characters and repeatedly merges the most frequent adjacent pair into a new token until it reaches a target vocabulary size. GPT-style models use byte-level BPE, which starts from raw bytes, so any script or emoji can be encoded.
  • WordPiece, used by BERT, chooses merges that most increase the likelihood of the training data. Continuation pieces carry a ## prefix, so "playing" can become "play" and "##ing."
  • Unigram language model starts from a large candidate vocabulary and prunes it down. It usually runs through the SentencePiece library, which treats input as a raw stream with spaces marked as "▁", so it needs no language-specific pre-splitting. That makes SentencePiece popular for multilingual models such as T5 and XLM-R.

Approach

"unbelievably" becomes (illustrative)

Vocabulary size

Main trade-off

Word

unbelievably

Open-ended, very large

Unseen words collapse to <UNK>

Character

u · n · b · e · l · i · e · v · a · b · l · y

A few hundred symbols

Very long sequences, weak meaning per token

Subword (BPE, WordPiece, Unigram)

un · believ · ably

Fixed, tens of thousands

Splits depend on the training corpus

Try it yourself

from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("bert-base-uncased")
pieces = tok.tokenize("Tokenization handles unbelievably rare words.")
ids = tok.convert_tokens_to_ids(pieces)
print(pieces)  # subword pieces, e.g. 'token', '##ization', ...
print(ids)     # the integers the model actually sees

Practical tips

  • Always use the tokenizer that ships with your pretrained model. Token ids from one vocabulary are meaningless to a model trained on another.
  • Token counts are not word counts. A common rule of thumb for English is roughly four characters, or three-quarters of a word, per token.
  • Languages that were rare in a tokenizer's training data, such as Urdu or Pashto, usually split into many more tokens for the same meaning. That raises API cost and shortens the effective context window.
  • For specialized domains like law, medicine or source code, consider training your own tokenizer with the Hugging Face tokenizers library.

Stop Words: When to Remove Them and When Not To

Stop words are very common words such as "the," "is," "at," "of" and "and." Removing them was standard practice in classic NLP, because they appear everywhere and carry little topical meaning on their own.

Why people remove them

In a bag-of-words model, stop words dominate the counts without helping to tell documents apart. Dropping them shrinks the feature space, speeds up training and makes keyword-driven tasks like topic modeling and clustering cleaner. Leave them in an LDA topic model, and your "topics" fill up with "the" and "and."

Why removal backfires

Standard stop-word lists contain words that matter a great deal for some tasks. NLTK's English list includes "not," "no," "nor," "against" and "don't." Strip them for sentiment analysis and "not good" turns into "good."

The damage goes further. Removing "who," "what" and "when" deletes the question type in question answering. Translation, summarization and text generation need every word to produce grammatical output. Hamlet's "to be or not to be" vanishes completely.

Guidelines

  • Remove them for topic modeling, keyword extraction, simple retrieval baselines and word clouds.
  • Keep them for sentiment analysis, intent detection, question answering, translation and anything using transformers or LLMs.
  • Customize the list. Take negations out, and add domain noise words instead. In product reviews, "product," "item" and "bought" may be more useless than "the."
  • Let the data decide. TF-IDF already down-weights frequent words, and scikit-learn's max_df=0.9 drops any term found in over 90% of documents, which acts as a data-driven stop list.
from nltk.corpus import stopwords

stops = set(stopwords.words("english"))
stops -= {"not", "no", "nor", "against"}  # keep negations
filtered = [t for t in tokens if t.lower() not in stops]

Stemming and Lemmatization: Reducing Words to Their Roots

Stemming and lemmatization both collapse forms like "run," "runs," "running" and "ran" into a shared base. They differ in how they do it, and that difference decides which one fits your task.

Stemming

A stemmer chops affixes off words with fixed rules and no dictionary. The Porter stemmer (1980) is the classic. The Snowball stemmer, also called Porter2, refines it and supports several languages, while the Lancaster stemmer is more aggressive.

Stemming is fast and needs no linguistic resources, but its output is often not a real word. It also makes two kinds of mistakes. Over-stemming merges unrelated words, so "universe" and "university" both become "univers." Under-stemming misses related ones, so "ran" never gets linked to "run."

Lemmatization

A lemmatizer returns the dictionary form, or lemma, using a vocabulary and grammar rules. It knows that "was" comes from "be," that "studies" comes from "study," and that the adjective "better" comes from "good."

Lemmatization needs to know each word's part of speech. "Meeting" is a noun in "the meeting ran late" and a verb in "we are meeting at noon." NLTK's WordNet lemmatizer assumes every word is a noun unless told otherwise, a common source of silent bugs, while spaCy tags parts of speech first and lemmatizes automatically.

Word

Porter stem

Lemma (with correct part of speech)

running

run

run

studies

studi

study

was

wa

be

better

better

good

university

univers

university

universe

univers

universe

organization

organ

organization

Which one should you use?

  • Stemming suits search engines and retrieval over large corpora, where speed matters and users never see the stems. Search tools like Elasticsearch ship stemming analyzers for exactly this reason.
  • Lemmatization suits tasks where output must be readable or meaning must stay precise. It is also the safer choice for morphologically rich languages such as Arabic, Turkish and Urdu, where crude suffix rules break down.
  • Neither is usually needed with transformer models. Subword tokens and contextual embeddings already link "run" and "running," and they keep the tense and number information that both techniques throw away.
import spacy
from nltk.stem import PorterStemmer

nlp = spacy.load("en_core_web_sm")
stemmer = PorterStemmer()

doc = nlp("The studies were better than earlier findings.")
for token in doc:
    print(f"{token.text:<10} stem={stemmer.stem(token.text):<8} lemma={token.lemma_}")

Adding Structure: POS Tagging, Named Entities and Parsing

Beyond splitting and normalizing, many pipelines annotate text with grammatical and semantic structure. These annotations become features for classic models, filters for rule-based systems and building blocks for information extraction.

Part-of-speech tagging

Part-of-speech (POS) tagging labels each token with its grammatical role: noun, verb, adjective, adverb and so on. Most tools use either the Universal Dependencies tag set (NOUN, VERB, ADJ, PROPN) or the finer Penn Treebank tags (NN, VBD, JJ).

POS tags help in many places. They pick the right lemma, pull out noun phrases as keywords, find the adjectives customers use about a product, and separate "book" the object from "book" the action. Good taggers reach around 97% accuracy on edited English news, though accuracy drops on social media and specialist text.

Named entity recognition

Named entity recognition (NER) finds spans that refer to real-world things and labels their type: PERSON, ORG (organization), GPE (countries and cities), DATE, MONEY and more. In "Ayesha Malik opened Acme Corp's new office in Lahore on Monday," a good model finds a person, an organization, a city and a date.

NER drives resume parsing, news analytics, support-ticket routing and redaction of personal data. Pretrained models handle general entities well, but domain entities such as drug names, legal clauses or product codes usually need custom labeled data. spaCy's training tools and Hugging Face token-classification models make that fine-tuning practical.

Dependency parsing

Dependency parsing finds grammatical relationships between words: which word is the subject of a verb, which is its object, and what each modifier attaches to. In "The battery of this phone drains fast," the parser links "drains" to its subject "battery" and "battery" to "phone."

That structure powers aspect-based sentiment analysis, where a system attaches the complaint "drains fast" to the battery rather than the phone as a whole. It also supports relation extraction (who acquired whom) and rule-based question answering.

Running it all at once

spaCy runs tokenization, tagging, parsing, lemmatization and NER in a single call:

import spacy

nlp = spacy.load("en_core_web_sm")
doc = nlp("Ayesha Malik opened Acme Corp's new office in Lahore on Monday.")

for token in doc:
    print(token.text, token.pos_, token.dep_, token.head.text)

for ent in doc.ents:
    print(ent.text, ent.label_)

Load only the components you need with spacy.load("en_core_web_sm", disable=["parser"]), and stream large datasets through nlp.pipe(texts, batch_size=256). Both changes can make processing several times faster on big corpora.

Do you still need this in the LLM era?

Less often than before, because end-to-end models learn much of this structure on their own. Explicit annotation still earns its place when you need interpretable rules, cheap processing at scale, structured output for a database, or features for a small dataset where deep models would overfit.

Turning Text into Numbers: Bag of Words, N-grams and TF-IDF

Models cannot read tokens; they need numbers. Representation, also called vectorization or feature extraction, maps each document to a vector. The classical methods are simple and fast, and they are surprisingly hard to beat on many tasks.

Bag of words

Bag of words (BoW) builds a vocabulary from every token in the training set and represents each document as a vector of counts. Take two short documents:

  1. "the food was good"
  2. "the food was not good, the service was slow"

With the vocabulary [the, food, was, good, not, service, slow], document 1 becomes [1, 1, 1, 1, 0, 0, 0] and document 2 becomes [2, 1, 2, 1, 1, 1, 1].

BoW ignores word order entirely, hence the name. The vectors are also sparse: with a 50,000-word vocabulary, a tweet has only a handful of non-zero entries. Sparse matrix formats keep that cheap to store and compute.

N-grams

N-grams restore some word order by treating runs of n tokens as features. Bigrams from document 2 include "not good" and "service was," so the model can finally tell "good" from "not good." Most practical setups use unigrams plus bigrams, since longer n-grams grow the vocabulary fast and rarely pay off without lots of data.

Character n-grams, such as every 3- to 5-character slice of a word, work well for noisy text, misspellings and language identification. "gooood" and "good" share many slices even though they are different words.

TF-IDF

Raw counts overvalue common words. TF-IDF (term frequency times inverse document frequency) weights each term by how often it appears in a document and how rare it is across the collection:

\text{tf-idf}(t, d) = \text{tf}(t, d) \times \log \frac{N}{\text{df}(t)}

Here N is the number of documents and df(t) is the number of documents containing term t. A word found in every document gets a weight near zero, while a word frequent in one document but rare elsewhere scores high. Libraries add smoothing; scikit-learn, for example, uses ln((1 + N) / (1 + df)) + 1 and scales each row to unit length, so its numbers differ slightly from the textbook formula.

from sklearn.feature_extraction.text import TfidfVectorizer

docs = ["the food was good", "the food was not good, the service was slow"]
vec = TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True)
X = vec.fit_transform(docs)          # sparse matrix: documents x features
print(X.shape)
print(vec.get_feature_names_out())   # includes bigrams like 'not good'

Why classical methods still matter

TF-IDF with logistic regression or a linear SVM trains in seconds on a laptop, is easy to explain, and makes a strong baseline for spam detection, topic classification and many sentiment tasks. Build one before reaching for a transformer. If the deep model beats it by only a point or two, the simpler system may be the better product.

The limit is meaning. To these vectors, "car" and "automobile" are as unrelated as "car" and "banana." Closing that gap is the job of embeddings.

Word Embeddings and Contextual Embeddings

Embeddings represent words as dense vectors, usually a few hundred to a few thousand numbers, where similar meanings sit close together. They replaced sparse count features as the default representation during the 2010s and remain central today.

The distributional idea

Embeddings rest on an old idea from linguist J. R. Firth: "You shall know a word by the company it keeps." Words that appear in similar contexts tend to mean similar things. "Coffee" and "tea" both show up near "cup," "drink" and "morning," so their vectors end up close together.

Static embeddings: Word2Vec, GloVe and fastText

  • Word2Vec (Google, 2013) trains a shallow neural network to predict a word from its neighbors (CBOW) or the neighbors from a word (skip-gram). Its vectors famously support analogies such as king − man + woman ≈ queen.
  • GloVe (Stanford, 2014) builds vectors from word co-occurrence counts gathered across the whole corpus.
  • fastText (Facebook AI Research, 2016) represents each word as a bag of character n-grams. It can therefore build vectors for unseen and misspelled words, which helps with noisy user text and morphologically rich languages.

These are called static embeddings because each word gets exactly one vector. "Bank" has the same vector in "river bank" and "bank loan," so two unrelated meanings get blended together.

Contextual embeddings

Contextual models produce a different vector for every occurrence of a word, based on its sentence. ELMo (2018) did this with bidirectional LSTMs. BERT (2018) did it with transformers and self-attention, and it reshaped the field: pretrain once on huge amounts of text, then fine-tune for many tasks.

In BERT, the two "banks" get distinct vectors, each shaped by its neighbors. That handles ambiguity, negation and long-range context far better than static vectors.

Sentence and document embeddings

Semantic search, clustering, duplicate detection and retrieval-augmented generation (RAG) all need one vector per sentence or document. Averaging word vectors is a weak baseline. Models trained for the job, such as Sentence-BERT and its successors in the sentence-transformers library, produce vectors whose cosine similarity tracks meaning.

from sentence_transformers import SentenceTransformer, util

model = SentenceTransformer("all-MiniLM-L6-v2")
a = model.encode("How do I reset my password?")
b = model.encode("I forgot my login credentials")
c = model.encode("What are your store hours?")
print(util.cos_sim(a, b), util.cos_sim(a, c))  # expect the first to be clearly higher

The first two questions share almost no words, so TF-IDF would call them unrelated. The embedding model recognizes that they ask nearly the same thing.

Choosing a representation

Representation

Best when

Watch out for

TF-IDF

Small data, need speed and interpretability, keyword-heavy tasks

No sense of synonyms or meaning

Static embeddings (fastText, GloVe)

Lightweight semantic features, low-resource languages, edge devices

One vector per word, so ambiguity is blurred

Contextual and sentence embeddings

Semantic search, RAG, meaning-heavy classification

Higher compute cost and model size

How Preprocessing Changes for Transformers and LLMs

With transformer models, the rule flips: do less. Pretrained models learned from natural, messy text, so the closer your input stays to that, the better they perform. Heavy classic preprocessing strips out exactly the cues they rely on.

What to stop doing

  • Removing stop words. "Not," "no" and "but" carry meaning that attention layers use.
  • Stemming and lemmatizing. The model already relates word forms, and it needs tense and number.
  • Stripping punctuation and case for cased models. Question marks, quotation marks and capitals are signals.
  • Building your own vocabulary. The model's tokenizer defines the vocabulary, and your job is to use it correctly.

What still matters

  • Cleaning. Remove HTML, boilerplate, encoding errors and duplicates. Junk tokens waste the context window and confuse the model.
  • The matching tokenizer and its special tokens. BERT expects [CLS] and [SEP]; instruction-tuned LLMs expect their chat template, applied with tokenizer.apply_chat_template. Hand-formatting prompts for a chat model is a classic cause of quietly worse output.
  • Maximum length. BERT-style models usually accept 512 tokens. Longer documents need truncation, a sliding window with overlap, or chunking.
  • Padding and attention masks. Sequences in a batch must share a length, so shorter ones get padded and the attention mask tells the model to ignore the padding.
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("distilbert-base-uncased")
batch = tok(
    ["Great service!", "The delivery was late and the box was damaged."],
    padding=True, truncation=True, max_length=128, return_tensors="pt",
)
print(batch["input_ids"].shape)   # (2, length of the longest sequence)
print(batch["attention_mask"])    # 1 = real token, 0 = padding

Chunking for retrieval and RAG

Retrieval-augmented generation splits documents into chunks, embeds each chunk, and hands the most relevant ones to an LLM. Chunking has become one of the most important preprocessing decisions. Chunks that are too small lose context, and chunks that are too large dilute relevance and waste tokens.

Sensible defaults help. Split on natural boundaries such as headings, paragraphs and sentences rather than fixed character counts. Keep a small overlap between neighboring chunks, attach metadata like the source title and section heading, and tune chunk size against measured retrieval quality.

Preprocessing for LLM training data

At the scale of LLM pretraining, text processing becomes data curation. Pipelines run language identification, quality filtering with heuristics and classifiers, large-scale deduplication, removal of personal information and filtering of toxic content. Open web datasets such as C4 and FineWeb were built with exactly these steps, because cleaner data gives a better model for the same compute.

Multilingual, Code-Mixed and Low-Resource Text

Most NLP tutorials assume clean English, and real-world text often is not. Billions of people write in other scripts, in languages with little labeled data, and in a mix of languages within a single sentence.

Scripts and word boundaries

  • No spaces between words. Chinese, Japanese and Thai break whitespace tokenization completely. They need dedicated segmenters, such as jieba for Chinese or MeCab for Japanese, or space-independent subword tokenizers like SentencePiece.
  • Arabic-script languages. Arabic, Persian, Urdu and Pashto run right to left, letters change shape by position, and short vowels are usually omitted. Look-alike letters also have separate code points: Arabic "ي" (U+064A) versus Persian and Urdu "ی" (U+06CC), or Arabic "ك" versus "ک". Map them to one form consistently, or your vocabulary quietly splits in two.
  • Invisible joiners. Urdu and Persian use the zero-width non-joiner inside words to control letter joining, which is why blanket removal of invisible characters can corrupt them.
  • Rich morphology. Turkish and Finnish pack many meanings into a single word, so word-level vocabularies explode and subword or morphological analysis pays off.

Code-mixing and transliteration

Across South Asia, Africa and many other regions, people switch languages mid-sentence and type their own language in Latin letters. "yaar meeting kal tak postpone kar do please" mixes Urdu and English in Roman Urdu. Spelling is not standardized either, so "kya," "kia" and "kyaa" are the same word.

A few practical steps help:

  • Detect language per token, not just per document.
  • Build normalization dictionaries for common Roman spellings from your own data.
  • Prefer character n-grams or subword models, which tolerate spelling variation.
  • Transliterate to the native script when good tools exist, or keep both versions as features.

Choosing models

Multilingual models such as mBERT, XLM-RoBERTa and multilingual sentence-transformers cover around 100 languages. They allow cross-lingual transfer: fine-tune on English labels, then apply the model to Urdu or Swahili with useful results. Quality is usually lower for languages that were scarce in pretraining data.

Check the tokenizer on your language before committing. If common words shatter into many tiny pieces, a model with better coverage, or one trained for your region, will likely serve you better.

Evaluate in the target language

Never assume English results transfer. Build a small, carefully labeled test set for each target language and dialect, including code-mixed samples. A few hundred well-chosen examples reveal more than any leaderboard.

Building a Complete Pipeline in Python

Here is a full, reusable pipeline for a review sentiment classifier that puts the earlier sections together. It wraps preprocessing in a scikit-learn transformer, so the exact same steps run during training and prediction.

The preprocessor

import re
import unicodedata

import emoji
import spacy
from bs4 import BeautifulSoup
from sklearn.base import BaseEstimator, TransformerMixin

URL_RE = re.compile(r"https?://\S+|www\.\S+")
MENTION_RE = re.compile(r"@\w+")
ELONGATED_RE = re.compile(r"([a-zA-Z])\1{2,}")  # letters only, so 1000 stays 1000
NEGATIONS = {"not", "no", "nor", "never", "n't"}
KEEP_PUNCT = {"!", "?"}

class TextPreprocessor(BaseEstimator, TransformerMixin):
    def __init__(self, lowercase=True, lemmatize=True, remove_stopwords=True):
        self.lowercase = lowercase
        self.lemmatize = lemmatize
        self.remove_stopwords = remove_stopwords

    def fit(self, X, y=None):
        return self

    def _nlp(self):
        if not hasattr(self, "nlp_"):
            self.nlp_ = spacy.load("en_core_web_sm", disable=["parser", "ner"])
        return self.nlp_

    def clean(self, text):
        text = BeautifulSoup(text, "html.parser").get_text(" ")
        text = unicodedata.normalize("NFKC", text)
        text = URL_RE.sub(" URL ", text)
        text = MENTION_RE.sub(" USER ", text)
        text = emoji.demojize(text, delimiters=(" ", " "))
        text = ELONGATED_RE.sub(r"\1\1", text)
        return re.sub(r"\s+", " ", text).strip()

    def transform(self, X):
        cleaned = [self.clean(t) for t in X]
        results = []
        for doc in self._nlp().pipe(cleaned, batch_size=256):
            tokens = []
            for tok in doc:
                if tok.is_space or (tok.is_punct and tok.text not in KEEP_PUNCT):
                    continue
                if self.remove_stopwords and tok.is_stop and tok.lower_ not in NEGATIONS:
                    continue
                word = tok.lemma_ if self.lemmatize else tok.text
                tokens.append(word.lower() if self.lowercase else word)
            results.append(" ".join(tokens))
        return results

    def __getstate__(self):
        state = self.__dict__.copy()
        state.pop("nlp_", None)  # spaCy reloads lazily after unpickling
        return state

Training and evaluation

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline

model = Pipeline([
    ("prep", TextPreprocessor()),
    ("tfidf", TfidfVectorizer(tokenizer=str.split, token_pattern=None, lowercase=False,
                              ngram_range=(1, 2), min_df=2, sublinear_tf=True)),
    ("clf", LogisticRegression(max_iter=1000, class_weight="balanced")),
])

# texts: list of review strings; labels: list of 0/1 sentiment labels
X_train, X_test, y_train, y_test = train_test_split(
    texts, labels, test_size=0.2, random_state=42, stratify=labels)

model.fit(X_train, y_train)
print(classification_report(y_test, model.predict(X_test)))

Why it is built this way

  • No train/serve skew. Preprocessing lives inside the pipeline, so saving it with joblib.dump(model, "sentiment.joblib") saves the cleaning rules too. Your web app's API runs exactly what training ran.
  • Every choice is testable. Lowercasing, lemmatization and stop-word removal are parameters, so GridSearchCV can compare them with a grid like {"prep__lemmatize": [True, False], "prep__remove_stopwords": [True, False]}.
  • Signals are protected. Negations survive stop-word removal, and "!" and "?" survive punctuation removal. The vectorizer splits on spaces, so it keeps exactly the tokens the preprocessor produced.
  • It is fast enough. spaCy runs with the parser and NER disabled and processes texts in batches.

The transformer alternative

For a transformer, keep only the clean step and let the model's own tokenizer do the rest:

from transformers import pipeline as hf_pipeline

sentiment = hf_pipeline("sentiment-analysis",
                        model="distilbert-base-uncased-finetuned-sst-2-english")
prep = TextPreprocessor()
reviews = ["<p>Not bad at all!!!</p>", "worst. service. ever."]
print(sentiment([prep.clean(r) for r in reviews]))

Run both on the same test set. The comparison tells you whether the transformer's extra accuracy is worth its extra cost for your product.

Common Mistakes and Best Practices

Most text-processing bugs do not crash anything. They quietly lower accuracy or inflate test scores, which makes them easy to miss and expensive to find later. These are the ones that show up again and again.

Mistake

What goes wrong

Better approach

Copying a tutorial's preprocessing recipe

Steps that suit one task damage another

Choose and test each step for your task

Removing all stop words for sentiment

"not good" turns into "good"

Keep negations, or skip removal entirely

Fitting the vectorizer on all the data

Test-set statistics leak into training and scores look inflated

Fit on the training split only, inside a pipeline

Different preprocessing in training and production

The model sees inputs unlike its training data, and accuracy drops silently

Ship preprocessing as part of the saved model

Stripping punctuation before handling URLs

"https://site.com" becomes "httpssitecom"

Replace structured patterns first

Heavy cleaning for transformer models

Removes cues the model learned to use

Remove noise only, and use the model's tokenizer

Ignoring encoding and Unicode

Mojibake, duplicate tokens, broken non-Latin text

Read as UTF-8, normalize with NFC or NFKC, repair with ftfy

Deduplicating after the train/test split

Near-copies on both sides inflate scores

Deduplicate before splitting

Lemmatizing without part-of-speech tags

WordNet treats every word as a noun

Pass POS tags, or use spaCy

Never reading the data

Template text, wrong languages and label noise go unnoticed

Read random samples at every stage

Habits that pay off

  • Start with a baseline. TF-IDF plus logistic regression gives you a number to beat in an afternoon.
  • Change one step at a time. Measure each preprocessing change on a fixed validation set, so you know what actually helped.
  • Version everything. Keep the raw text, and version preprocessing code alongside the model it fed.
  • Unit-test the pipeline. Feed it empty strings, emojis, URLs, mixed scripts and very long documents, and check the outputs.
  • Look at outputs, not just metrics. Before-and-after pairs of real examples catch bugs that an accuracy score hides.
  • Protect privacy early. Mask emails, phone numbers and ID numbers at the start, so they never reach logs or training sets.

Conclusion: Make Every Step a Deliberate Choice

Text processing is where NLP projects are quietly won or lost. The algorithms are rarely the hard part. The hard part is understanding your text well enough to know which steps preserve its meaning and which ones destroy it.

A few principles carry across every project:

  1. Look at your data before and after every step.
  2. Match the pipeline to the model: rich preprocessing for classical models, light cleaning for transformers.
  3. Protect the signals your task depends on, whether that is negation, punctuation, case, numbers or emojis.
  4. Package preprocessing with the model, so training and production never drift apart.
  5. Measure every change against a strong, simple baseline.

Where to go next

Pick a small, real dataset, such as product reviews, support tickets or social posts in the languages you work with, and build the pipeline from this guide end to end. Then experiment: turn lemmatization off, keep stop words, swap TF-IDF for sentence embeddings, and watch how the metrics move. One afternoon of that teaches more than a dozen tutorials.

Good next topics are fine-tuning transformers with Hugging Face, building semantic search or a RAG system on sentence embeddings, and training a custom NER model with spaCy. Helpful free resources include the Hugging Face course, the spaCy documentation and its online course, and Jurafsky and Martin's textbook Speech and Language Processing, whose third-edition draft is available online.

Language models are spreading into every kind of product. The developers who understand what happens to text before it reaches a model will be the ones who make those products work reliably.

Comments

Popular posts from this blog

Multimodal AI Explained: How Text, Image & Voice Merge Into One Model

  The Rise of Multimodal AI: Text, Image, and Voice in One Model For most of the last decade, AI systems were narrow specialists: a language model that only understood text, an image classifier that only understood pictures, a speech recognition system that only understood audio. Getting these systems to work together meant stitching together separate pipelines, converting between formats at every handoff, and accepting the errors and awkwardness that came with each conversion step. That era is ending. Modern multimodal AI systems process text, images, audio, and increasingly video within a single unified model, reasoning across all of them together rather than treating each as a separate problem solved by a separate system. This guide explains what multimodal AI actually is, the real architectural difference between "bolted-together" and "natively unified" multimodal systems, where this technology is already changing real products, and an honest look at where it ...

PyTorch Explained: The Complete Guide to Deep Learning & Neural Networks in 2026

  PyTorch: The Complete Guide to Deep Learning's Most Popular Framework Introduction If you've trained a neural network, fine-tuned a language model, or experimented with a diffusion-based image generator in the last several years, there's a strong chance PyTorch was somewhere underneath it. Originally released by Facebook AI Research (now Meta AI) in 2016, PyTorch has grown from a research-focused alternative to established frameworks into the dominant tool in the deep learning world — powering everything from academic papers to some of the largest AI systems ever deployed in production. This guide takes a deep, practical look at PyTorch: what it is, why it was designed the way it was, how its core components fit together, and how to actually use it to build, train, and deploy real models. Whether you're completely new to deep learning or you've used other frameworks and want to understand what makes PyTorch different, this article will walk you through everythi...

AI Job Displacement 2026: What the Data Really Shows

  AI and Job Displacement: What's Actually Happening in 2026 Few questions about AI generate more anxiety, and more contradictory headlines, than what it's actually doing to jobs. One week brings a report of tens of thousands of layoffs attributed to AI; the next brings a forecast of net job creation once new AI-related roles are counted. Both can be true at once, describing different parts of a genuinely uneven, still-unfolding transition. This guide sets aside both the most alarmist and the most dismissive framings and works through what the actual 2026 data — from government labor statistics, corporate layoff tracking, and major research institutions — shows about where AI is displacing work, where it's mainly changing hiring rather than firing, and where the picture remains genuinely uncertain. Given how fast this data changes, treat the specific figures here as a snapshot of 2026, not a permanent verdict. 1. The Honest Headline: Displacement Is Real, Concentrated, ...