Skip to main content

35 Best Free AI Tools to Make Your Work Easier in 2026

The Practical 2026 Guide to Free AI for Writing, Research, Design, Video, Meetings and Code Updated October 2026 · 39 min read Table of Contents Introduction: Work Smarter, Not Harder, With Free AI How We Chose These 35 Tools How to Choose the Right AI Tool for Your Work Category 1: AI Assistants for Everyday Work Category 2: Research and Learning Category 3: Writing and Language Category 4: Design and Visuals Category 5: Presentations Category 6: Video and Audio Category 7: Meetings and Automation Category 8: Coding and App Building Category 9: Local and Private AI All 35 Tools at a Glance Ready-Made Workflows: Combining Free Tools Using Free AI Tools Safely and Responsibly Five Prompting Habits That Make Every Tool Better Frequently Asked Questions Conclusion: Start Small, Save Hours Introduction: Work Smarter, Not Harder, With Free AI A few years ago, using artificial intelligence at work meant hiring data scientists or paying for expensive enterprise software. In 2026, some of the...

Pandas for NLP Text Processing: Complete Python Guide (2026)



The Complete Python Guide to Cleaning, Exploring and Modelling Text Data with pandas 3.x

Updated October 2026 · 33 min read

Table of Contents

  1. Introduction: Why Pandas Still Matters for NLP in 2026
  2. What Changed in Pandas 3.0 for Text Data
  3. Setting Up Your NLP Environment
  4. Loading Text Data the Right Way
  5. Exploratory Analysis of Text Data
  6. Mastering the .str Accessor
  7. Regular Expressions for Text Cleaning
  8. Unicode Normalisation and Multilingual Text
  9. Building a Reusable Cleaning Pipeline
  10. Tokenization: Turning Sentences Into Pieces
  11. Stopwords, Stemming and Lemmatization
  12. Word Frequencies and N-grams with explode()
  13. Feature Engineering: From Text to Numbers
  14. End-to-End Project: A Sentiment Classifier
  15. Enriching Rows: Sentiment Scores and Named Entities
  16. Embeddings and Semantic Search in a DataFrame
  17. Performance: Processing Millions of Rows
  18. Best Practices and Common Mistakes
  19. Pandas NLP Cheat Sheet
  20. Frequently Asked Questions
  21. Conclusion: Your Text Workbench Is Ready

Introduction: Why Pandas Still Matters for NLP in 2026

Every Natural Language Processing project, no matter how advanced, starts in the same humble place: a pile of messy text. Customer reviews full of typos and emojis. Support tickets with copied email signatures. Tweets stuffed with hashtags, mentions and links. Scraped web pages that still carry stray HTML tags. Before any transformer model, large language model or vector database can do something clever with that text, someone has to load it, inspect it, clean it, organise it and turn it into features. In the Python world, that “someone” is very often pandas.

It is easy to assume that pandas is old news in an era of giant language models. The opposite is true. Large models have made the data side of NLP more important, not less. Fine-tuning a model on 50,000 dirty examples produces a dirty model. Evaluating a chatbot means comparing thousands of generated answers against references. Building a retrieval system means chunking, deduplicating and labelling documents. All of those jobs are tabular at heart: one row per document, one column per property. That is exactly the shape pandas was built for.

Think of pandas as the workbench of an NLP workshop. The power tools (spaCy, scikit-learn, Hugging Face Transformers, sentence-transformers) do the heavy cutting, but the workbench is where you lay out your material, measure it, mark it, hold it steady and check the result. A good workbench does not replace the tools. It makes every tool easier and safer to use.

In this guide you will learn how to use pandas as the backbone of a complete text-processing workflow. We will move step by step, from loading raw text, through exploratory analysis and cleaning, to tokenization, feature engineering, classification, entity extraction, embeddings and performance tuning for large datasets. Every section includes working Python code written for pandas 3.x, the major release that changed how pandas stores and handles strings.

Who This Guide Is For

This guide is written for three kinds of readers:

  • Beginners in data science who know basic Python and want a practical, structured path into text data.
  • Developers (web, backend, full-stack) who need to analyse user-generated text such as reviews, comments or tickets without becoming full-time ML researchers.
  • Students and practitioners preparing datasets for machine learning models, including fine-tuning and evaluation of modern language models.

You do not need prior NLP experience. If you can write a Python function and you have seen a DataFrame before, you are ready.

What You Will Build

By the end, you will have a reusable, production-style text pipeline that can:

  1. Load text from CSV, JSON Lines and Parquet files, including very large files.
  2. Profile a text dataset in minutes (lengths, duplicates, empty rows, class balance, vocabulary).
  3. Clean text with vectorized string methods and regular expressions.
  4. Tokenize, remove stopwords, and lemmatize at scale.
  5. Count words and n-grams with the elegant explode() pattern.
  6. Engineer handcrafted and TF-IDF features and train a text classifier.
  7. Enrich rows with sentiment, named entities and semantic embeddings.
  8. Process millions of rows without running out of memory.

How to read this guide

Each section builds on the previous one, but every section also stands alone. If you already know the basics, jump straight to the pipeline, feature engineering or performance sections using the table of contents.

What Changed in Pandas 3.0 for Text Data

Pandas 3.0.0 was released in January 2026 after almost three years of development. It is the first major version since 2.0, and two of its headline changes directly affect anyone working with text. If you learned pandas from older tutorials, this section will save you hours of confusion.

1. Strings Finally Have Their Own Data Type

For most of its history, pandas stored text in columns with the NumPy object dtype. An object column is like a box that can hold anything: strings, numbers, lists, dictionaries, even other DataFrames. That flexibility came at a cost. Operations were slow, memory usage was high, and you could never be sure what was inside a column labelled object.

In pandas 3.0, text columns are automatically inferred as a dedicated str dtype. When the PyArrow library is installed, this dtype is backed by Apache Arrow, which stores strings contiguously in memory. The practical effect is that common string operations are noticeably faster and use less memory, and a column’s type now tells you honestly that it contains text.

Python
import pandas as pd

reviews = pd.Series(["Great product!", "Terrible support", None])
print(reviews.dtype)
# str <- pandas 3.x
# object <- pandas 2.x and earlier

Missing values in a str column are represented as NaN, and the column only accepts strings or missing values. If you try to put an integer inside a str column, pandas will complain instead of silently mixing types. For NLP work this strictness is a gift: the classic bug where a review column secretly contains a few floats (usually NaNs from empty cells) and crashes your tokenizer becomes much easier to catch.

Code that checks for object dtype

Older code often detects text columns with df.select_dtypes(include="object") or if df[col].dtype == object. In pandas 3.x those checks miss the new str columns. Use df.select_dtypes(include=["object", "str"]) or pd.api.types.is_string_dtype(df[col]) instead.

2. Copy-on-Write Is Now the Only Mode

The second big change is Copy-on-Write (CoW). In older pandas, selecting a subset of a DataFrame sometimes returned a view (sharing memory with the original) and sometimes a copy. You could not easily tell which, so pandas threw the infamous SettingWithCopyWarning whenever you modified a selection.

In pandas 3.0, every selection behaves as if it were a copy. Pandas still shares memory under the hood for speed, but the moment you modify something, it makes a real copy. This has two consequences for text processing:

  • Chained assignment no longer works. Code like df["text"][df["text"].isna()] = "" silently does nothing to df. Use df.loc[mask, "text"] = "" or, even better, return new columns with assign().
  • Defensive .copy() calls are no longer needed just to silence warnings, and SettingWithCopyWarning is gone.
Python
# Wrong in pandas 3.x: modifies a temporary copy, not df
df["text"][df["text"].isna()] = ""

# Right: one indexing operation on the DataFrame itself
df.loc[df["text"].isna(), "text"] = ""

# Also right, and more idiomatic: build a new column
df = df.assign(text=df["text"].fillna(""))

3. A Cleaner Way to Write Column Expressions with pd.col

Pandas 3.0 also introduced initial support for pd.col(), an expression syntax that lets you refer to a column by name inside assign() without writing a lambda. It reads beautifully in text pipelines:

Python
df = df.assign(
n_chars=pd.col("text").str.len(),
text_lower=pd.col("text").str.lower(),
)

Under the hood this is equivalent to lambda d: d["text"].str.len(), but it is shorter and easier to scan. Because the syntax is new, many tutorials and Stack Overflow answers still use lambdas, and both styles work. In this guide we use whichever is clearer for the example at hand.

4. Other Changes Worth Knowing

  • Pandas 3.x requires Python 3.11 or newer.
  • str.replace() treats the pattern as a literal string by default (this changed back in 2.0, but many people still miss it). Pass regex=True whenever you use a regular expression.
  • Datetime columns now default to an inferred resolution (often microseconds) instead of always using nanoseconds, which matters if you analyse timestamps of social media posts.
  • The upgrade path the pandas team recommends is to move to pandas 2.3 first, fix all warnings, and only then upgrade to 3.0.
TopicPandas 2.x habitPandas 3.x habit
Text column typeobjectstr (Arrow-backed when PyArrow is installed)
Detecting text columnsdtype == objectpd.api.types.is_string_dtype()
Modifying a subsetChained indexing + .copy()Single .loc[] call or assign()
Regex replaceSometimes implicitAlways regex=True
New column from expressionlambda d: ...pd.col("name")... or lambda

Setting Up Your NLP Environment

A clean environment prevents most “it works on my machine” problems. We recommend creating a virtual environment per project. The examples below use venv, but uv, conda or Poetry work equally well.

Terminal
python -m venv nlp-env
source nlp-env/bin/activate # Windows: nlp-env\Scripts\activate

pip install --upgrade "pandas>=3.0" pyarrow
pip install scikit-learn nltk spacy
pip install sentence-transformers matplotlib
python -m spacy download en_core_web_sm

Here is what each package does in our workflow:

PackageRole in the pipeline
pandasLoad, inspect, clean, reshape and summarise text tables
pyarrowFast, memory-efficient string storage and Parquet support
scikit-learnTF-IDF vectorization, classification, evaluation
nltkStopword lists, VADER sentiment, classic tokenizers
spacyIndustrial-strength tokenization, lemmatization and NER
sentence-transformersSemantic embeddings for similarity and search
matplotlibQuick charts of distributions and frequencies

NLTK needs a one-time download of its data files:

Python
import nltk
nltk.download("stopwords")
nltk.download("vader_lexicon")

Pin your versions

Text pipelines are sensitive to library versions: a new tokenizer release can change token counts, and a new spaCy model can change lemmas. Save your environment with pip freeze > requirements.txt so results are reproducible months later.

Loading Text Data the Right Way

Text data arrives in many shapes. The goal of this step is simple: get every document into a DataFrame with one row per document and a clearly named text column. Everything else in this guide depends on that shape.

Reading CSV Files Safely

CSV is the most common format for text datasets and also the most fragile, because text itself contains commas, quotes and line breaks. Pandas handles quoted fields well, but a few parameters make loading more robust:

Python
df = pd.read_csv(
"reviews.csv",
encoding="utf-8", # be explicit; try "utf-8-sig" for Excel exports
on_bad_lines="warn", # report malformed rows instead of crashing
dtype={"review_id": "str"},# keep IDs as text, never as numbers
)
print(df.shape)
print(df.dtypes)

If a file was exported from Excel on Windows, you may see a strange  character at the start of the first column name. That is a byte-order mark. Reading with encoding="utf-8-sig" removes it. If you see UnicodeDecodeError, the file is probably not UTF-8; try encoding="cp1252" or encoding="latin-1", then normalise everything to UTF-8 when you save.

JSON Lines: The Favourite Format of NLP Datasets

Many NLP datasets, chat logs and API exports use JSON Lines (.jsonl), where each line is a complete JSON object. Pandas reads them with one argument:

Python
df = pd.read_json("conversations.jsonl", lines=True)

When records are nested (for example a user object inside each message), flatten them with pd.json_normalize():

Python
import json

with open("tickets.jsonl", encoding="utf-8") as f:
records = [json.loads(line) for line in f]

df = pd.json_normalize(records, sep="_")
# columns like: id, body, customer_name, customer_country, tags

Parquet: The Best Format for Processed Text

Once you have cleaned a dataset, save it as Parquet rather than CSV. Parquet is a compressed, columnar format that preserves data types, loads much faster and takes a fraction of the disk space. It is also the format used by the Hugging Face Hub for datasets, so it fits naturally into modern NLP workflows.

Python
df.to_parquet("reviews_clean.parquet", index=False)
df = pd.read_parquet("reviews_clean.parquet", columns=["review_id", "text", "label"])

The columns= argument is a hidden superpower: Parquet only reads the columns you request, so loading a single text column from a wide table is fast.

Reading Files That Do Not Fit in Memory

If a CSV file is several gigabytes, reading it all at once may crash your notebook. Use chunksize to process it piece by piece:

Python
chunks = pd.read_csv("big_reviews.csv", chunksize=100_000)

total_words = 0
for chunk in chunks:
total_words += chunk["text"].str.split().str.len().sum()

print(f"Total words: {total_words:,}")

We will come back to large-scale processing in the performance section.

Our Example Dataset

To keep examples concrete, we will use a small product-review dataset throughout this guide. You can replace it with your own data at any point.

Python
data = {
"review_id": ["r1", "r2", "r3", "r4", "r5", "r6", "r7", "r8"],
"text": [
"Absolutely LOVE this phone!!! Battery lasts 2 days 😍 #happy",
"Terrible support. Emailed help@shop.com three times, no reply...",
"<p>Good value for money. Screen is bright.</p>",
"Worst purchase ever. Broke after 1 week!! https://bit.ly/x1",
None,
"Decent camera, average battery. Would buy again? Maybe.",
"Absolutely LOVE this phone!!! Battery lasts 2 days 😍 #happy",
" Delivery was FAST and packaging was great ",
],
"rating": [5, 1, 4, 1, 3, 3, 5, 5],
}
df = pd.DataFrame(data)
df["label"] = (df["rating"] >= 4).map({True: "positive", False: "negative"})

Notice that this tiny dataset already contains almost every problem you will meet in real life: uppercase shouting, repeated punctuation, emojis, hashtags, an email address, HTML tags, a URL, a missing value, an exact duplicate and extra whitespace. Real datasets simply have more of the same.

Exploratory Analysis of Text Data

Before you clean anything, look at your data. Exploratory Data Analysis (EDA) for text answers a handful of questions that decide every later step: How long are the documents? How many are empty or duplicated? Are the classes balanced? What does the vocabulary look like? Skipping EDA is like cooking without tasting.

Missing and Empty Text

A missing value (NaN) and an empty or whitespace-only string are different things, and both are common.

Python
n_missing = df["text"].isna().sum()
n_blank = df["text"].str.strip().eq("").sum()
print(f"Missing: {n_missing}, Blank: {n_blank}")

Decide early what to do with these rows. For training a classifier, you normally drop them. For an analytics dashboard, you might keep them and report them separately.

Document Length Statistics

Length is the single most informative property of a text dataset. It tells you how much context a model needs, how to choose a maximum sequence length, and whether some rows are junk.

Python
df = df.assign(
n_chars=pd.col("text").str.len(),
n_words=pd.col("text").str.split().str.len(),
)
print(df[["n_chars", "n_words"]].describe())

The describe() output gives you the mean, median (50%), and extreme values. Pay special attention to the maximum: a single 50,000-word “review” is usually a scraping error or a pasted document, and it can dominate memory and training time. A histogram makes the distribution visible:

Python
ax = df["n_words"].plot.hist(bins=30, title="Words per review")
ax.set_xlabel("Number of words")

Choosing a max length for transformer models

When you later fine-tune a transformer, a common rule of thumb is to choose a maximum token length that covers about 95% of your documents: df["n_words"].quantile(0.95). Remember that tokens are usually more numerous than words, often around 1.3 tokens per English word, and more for other languages.

Duplicates

Duplicate text inflates your dataset, biases models towards repeated examples and, worst of all, leaks examples between training and test sets, making your model look better than it is.

Python
exact_dupes = df.duplicated(subset="text", keep="first")
print(f"Exact duplicates: {exact_dupes.sum()}")

# Near-duplicates: compare a normalised version of the text
# (fine for English; for other scripts see "The Regex Engine Trap" below)
norm = df["text"].str.lower().str.replace(r"\W+", " ", regex=True).str.strip()
print(f"Normalised duplicates: {norm.duplicated().sum()}")

The normalised comparison catches rows that differ only by capitalisation, punctuation or spacing. Later, in the embeddings section, we will catch semantic duplicates that use different words to say the same thing.

Class Balance

If you plan to train a classifier, check how your labels are distributed:

Python
print(df["label"].value_counts(normalize=True).round(2))

A dataset with 95% positive reviews will produce a model that predicts “positive” for almost everything and still reports 95% accuracy. Knowing the balance in advance tells you to use better metrics (precision, recall, F1) and techniques such as class weights or stratified splitting.

A Quick Look at Vocabulary

Even before proper cleaning, a rough word count reveals the character of your data:

Python
top_words = (
df["text"].dropna()
.str.lower()
.str.findall(r"[a-z']+")
.explode()
.value_counts()
.head(15)
)
print(top_words)

At this stage the top words will be dominated by “the”, “and”, “was” and similar function words. That is expected, and it is exactly why we will remove stopwords later. You may also spot domain-specific vocabulary, brand names or recurring junk tokens (like “amp” from badly decoded &amp;) that you should handle during cleaning.

Build a One-Line Profile Function

Wrap the checks above into a function you can run on any new text dataset:

Python
def profile_text(df: pd.DataFrame, col: str = "text") -> pd.Series:
s = df[col]
words = s.str.split().str.len()
return pd.Series({
"rows": len(s),
"missing": s.isna().sum(),
"blank": s.str.strip().eq("").sum(),
"exact_duplicates": s.duplicated().sum(),
"median_words": words.median(),
"p95_words": words.quantile(0.95),
"max_words": words.max(),
})

print(profile_text(df))

Running this function the moment you receive a dataset gives you a five-second health check, and it often reveals problems that would otherwise surface only after hours of debugging a model.

Mastering the .str Accessor

If pandas is the workbench of NLP, the .str accessor is its toolbox. Every Series of text has a .str attribute that exposes dozens of string methods. These methods are vectorized: you call them once on the whole column, and pandas applies them to every row, automatically skipping missing values. No for loops, no apply, no try/except for None.

The mental model is simple. Anything you would do to a single Python string, you can usually do to a whole column by inserting .str in the middle:

Python
"Hello World".lower()          # one string
df["text"].str.lower() # every row in the column

The Essential Methods, Grouped by Purpose

The table below lists the methods you will use most often in NLP work. You do not need to memorise them; skim the list so you know what exists.

PurposeMethodExample
Caselower(), upper(), casefold(), title()s.str.casefold()
Whitespacestrip(), lstrip(), rstrip()s.str.strip()
Searchcontains(), startswith(), endswith(), match(), fullmatch()s.str.contains("refund", case=False)
Countlen(), count()s.str.count(r"!")
Replacereplace(), removeprefix(), removesuffix(), translate()s.str.replace(r"\d+", "<NUM>", regex=True)
Split & joinsplit(), rsplit(), partition(), join(), cat()s.str.split()
Extractextract(), extractall(), findall()s.str.findall(r"#\w+")
Sliceslice(), [ ], get()s.str[:100]
Test typeisalpha(), isdigit(), isspace(), isnumeric()s.str.isdigit()
Unicodenormalize(), encode(), decode()s.str.normalize("NFKC")
Dummiesget_dummies()tags.str.get_dummies(sep="|")

Case Normalisation: lower() vs casefold()

Most pipelines lowercase text so that “Battery”, “BATTERY” and “battery” count as the same word. lower() works for English, but casefold() is more aggressive and handles special cases in other languages (for example, the German “ß” becomes “ss”). For multilingual data, prefer casefold().

Be careful, though: case can carry meaning. “US” (the country) and “us” (the pronoun), or a customer shouting in ALL CAPS, are signals that a sentiment model might want. A good practice is to extract case-based features first, then lowercase.

Searching Text with contains()

contains() returns a Boolean mask, which makes it perfect for filtering:

Python
refund_mask = df["text"].str.contains(r"refund|money back|return", case=False, na=False)
refund_reviews = df[refund_mask]

Two arguments deserve attention. case=False makes the search case-insensitive, and na=False treats missing text as “does not match”. Without na=False, missing rows return NaN, and using a mask with NaN values for filtering raises an error.

Counting Patterns

Counting is the simplest form of feature engineering. The number of exclamation marks, question marks or uppercase words is often surprisingly predictive of sentiment or urgency:

Python
df = df.assign(
n_exclaim=pd.col("text").str.count("!"),
n_question=pd.col("text").str.count(r"\?"),
n_caps_words=pd.col("text").str.count(r"\b[A-Z]{2,}\b"),
)

Note that count() always interprets its argument as a regular expression, which is why the question mark is escaped as \?.

Extracting Structured Information

extract() pulls out the first match of a regex capture group into new columns, while findall() returns every match as a list. Together they turn unstructured text into structured data:

Python
# Every hashtag and mention, as lists
df["hashtags"] = df["text"].str.findall(r"#(\w+)")
df["mentions"] = df["text"].str.findall(r"@(\w+)")

# The first number followed by a time unit, split into two columns
durations = df["text"].str.extract(r"(?P<amount>\d+)\s*(?P<unit>day|week|month)s?")
df = df.join(durations)

Named groups like (?P<amount>...) become column names automatically, which keeps your code self-documenting.

Splitting Into Columns or Lists

split() returns a list per row by default. With expand=True it spreads the pieces into separate columns, which is handy for semi-structured text like “Category > Subcategory > Product”:

Python
paths = pd.Series(["Electronics > Phones > Android", "Home > Kitchen > Kettles"])
paths.str.split(" > ", expand=True)
# 0 1 2
# 0 Electronics Phones Android
# 1 Home Kitchen Kettles

Turning Tags Into Features with get_dummies()

Many datasets store multiple labels in one field separated by a delimiter. str.get_dummies() converts them into one-hot columns in a single call, which is ideal for multi-label classification:

Python
tags = pd.Series(["billing|refund", "shipping", "refund|shipping|damaged"])
tags.str.get_dummies(sep="|")
# billing damaged refund shipping
# 0 1 0 1 0
# 1 0 0 0 1
# 2 0 1 1 1

Regular Expressions for Text Cleaning

Regular expressions (regex) are a small language for describing patterns in text. They look intimidating at first, but in NLP you will reuse the same dozen patterns again and again. Learning them is one of the highest-return investments you can make.

A Two-Minute Regex Primer

PatternMeaningMatches
\dany digit7 in “7 days”
\wword character (letter, digit, underscore)a, Z, 5, _
\swhitespace (space, tab, newline)the gaps between words
.any character except newlineanything
+one or more of the previous item\d+ matches 2026
*zero or morea* matches “”, “a”, “aaa”
?optional (zero or one)colou?r matches both spellings
[...]any one character from a set[aeiou] matches a vowel
[^...]any character not in the set[^a-z ] matches symbols
\bword boundary\bcat\b does not match “concatenate”
( )capture group(\d+) days captures the number
|ORrefund|return

Always write regex patterns as raw strings (r"...") in Python so backslashes are not interpreted twice.

The Pattern Library You Will Reuse Forever

Here is a compact library of cleaning patterns for typical web and social text. Store it in a module (for example patterns.py) and import it into every project.

Python
import re

PATTERNS = {
"url": r"https?://\S+|www\.\S+",
"email": r"[\w.+-]+@[\w-]+\.[\w.-]+",
"mention": r"@\w+",
"hashtag": r"#(\w+)",
"html": r"<[^>]+>",
"number": r"\b\d+(?:[.,]\d+)?\b",
"repeat_punct": r"([!?.])[!?.]+",
"multi_space": r"\s+",
}

Applying Patterns Step by Step

Let us clean the example reviews one problem at a time. Watching each step makes the effect of every pattern obvious.

Python
s = df["text"].fillna("")

s = s.str.replace(PATTERNS["html"], " ", regex=True) # strip HTML tags
s = s.str.replace(PATTERNS["url"], " <URL> ", regex=True) # mask links
s = s.str.replace(PATTERNS["email"], " <EMAIL> ", regex=True) # mask emails
s = s.str.replace(PATTERNS["hashtag"], r"\1", regex=True) # keep hashtag word
s = s.str.replace(PATTERNS["repeat_punct"], r"\1", regex=True)# "!!!" -> "!"
s = s.str.replace(PATTERNS["multi_space"], " ", regex=True).str.strip()

Notice a few deliberate choices:

  • URLs and emails are replaced with placeholder tokens (<URL>, <EMAIL>) instead of being deleted. The fact that a review contains a link can be informative (spam often does), and placeholders also protect personal data.
  • Hashtags keep their word. #happy becomes happy, preserving the sentiment signal while removing the symbol.
  • Runs of punctuation are collapsed, not removed. “!!!” becomes “!”, so one exclamation mark still signals emphasis.
  • Whitespace is normalised last, because earlier replacements introduce extra spaces.

Order matters

Remove HTML before matching URLs (tags often contain links), and mask emails before removing symbols like @ (otherwise an email turns into two meaningless fragments). When a cleaning step behaves strangely, the cause is almost always the order of operations.

Handling HTML Entities and Emojis

Scraped text often contains HTML entities such as &amp;, &quot; or &#39;. Python’s standard library decodes them:

Python
import html
s = s.map(html.unescape)

Emojis are trickier. For sentiment analysis they are gold: 😍 and 😡 say more than many words. For a topic model they are noise. You have three options: keep them, remove them, or convert them into words. The third option keeps the meaning while making the text friendlier for classic models. The emoji package does this conversion (emoji.demojize("😍") returns ":smiling_face_with_heart-eyes:"). If you simply want to remove all emojis and other pictographs, a Unicode-range pattern works:

Python
import re

EMOJI = r"[\U0001F300-\U0001FAFF\U00002600-\U000027BF\U0001F1E6-\U0001F1FF]"
s_no_emoji = s.str.replace(EMOJI, "", regex=True, flags=re.UNICODE)

The flags=re.UNICODE argument looks redundant, but it matters in pandas 3.x. The next section explains why.

Unicode Normalisation and Multilingual Text

Text that looks identical can be stored as different sequences of bytes. The letter “é” can be one character or two (an “e” followed by a combining accent). Full-width characters such as “ABC” from East Asian keyboards look like normal letters but are different code points. Fancy quotes (“ ”) and plain quotes (“) are different characters too. If you do not normalise, your duplicate detection and word counts quietly go wrong.

The str.normalize() method applies the standard Unicode normalisation forms:

Python
s = s.str.normalize("NFKC")

NFKC is the most useful form for NLP. It composes accented characters into single code points and converts “compatibility” characters (full-width letters, ligatures like “fi”, superscripts) into their plain equivalents.

Working with Non-English Scripts

Pandas handles any language that Unicode supports, including Urdu, Arabic, Hindi, Chinese and many more. A few practical tips make multilingual work smoother:

  • Do not use [a-z] patterns for non-Latin text; they will delete everything. Use Unicode-aware classes instead (see the warning below).
  • Arabic-script languages like Urdu and Arabic contain optional diacritics and variant letter forms. Normalising those (for example, removing diacritics in the range \u064B to \u065F) often improves matching.
  • Detect the language first if your dataset mixes languages, and route each language to the right tokenizer and stopword list. Libraries such as langdetect, lingua or fastText’s language identification model can populate a lang column, after which groupby("lang") lets you process each language separately.
  • Roman Urdu and code-mixed text (Urdu written in Latin letters, often mixed with English) has no standard spelling, so normalising common variants with a small replacement dictionary can noticeably improve results.
Python
urdu = pd.Series([
"یہ فون بہت اچھا ہے",
"سروس بہت خراب تھی",
])
print(urdu.str.split().str.len()) # [5, 4]: word counts work out of the box

The Regex Engine Trap in Pandas 3.x

This is one of the most important and least documented details for multilingual NLP in pandas 3.x, so it deserves its own section.

When a column uses the Arrow-backed str dtype, pandas runs several regex methods (including str.replace(), str.contains() and str.count()) through Arrow’s compute engine, which uses Google’s RE2 regular-expression library instead of Python’s built-in re module. RE2 is fast and safe, but it differs from Python in two ways that can silently break text pipelines:

  1. \w, \d, \s and \b are ASCII-only in RE2. Letters such as “é”, “ß” or any Urdu, Arabic, Hindi or Chinese character do not count as word characters.
  2. Backreferences inside the pattern (like \1) and \U escapes are not supported and raise an ArrowInvalid error.

The first difference is the dangerous one, because it does not raise an error. Look at what happens to a few accented words:

Python
import re

words = pd.Series(["café naïve résumé"])

words.str.replace(r"\W+", "|", regex=True)
# 0 caf|na|ve|r|sum| <- accented letters treated as "non-word" and destroyed

words.str.replace(r"\W+", "|", regex=True, flags=re.UNICODE)
# 0 café|naïve|résumé <- correct: pandas falls back to Python's re module

words.str.replace(r"[^\p{L}\p{N}_]+", "|", regex=True)
# 0 café|naïve|résumé <- correct: explicit Unicode classes understood by RE2

For Urdu, Arabic or Hindi text the damage is even worse. Every letter in those scripts counts as “non-word” to RE2, so the first version replaces an entire Urdu sentence with a single | and your text is simply gone. Both corrected versions keep every Urdu word intact.

Meanwhile, str.findall() and str.extract() currently use Python’s re module, so \w is Unicode-aware there. That inconsistency is exactly why bugs slip through: one step of your pipeline handles Urdu correctly and the next one quietly destroys it.

Rules for regex on non-English text in pandas 3.x

Use explicit Unicode property classes such as \p{L} (any letter) and \p{N} (any number) with str.replace(), str.contains() and str.count(), or pass flags=re.UNICODE to make pandas use Python’s re engine (slower, but fully Unicode-aware and supporting backreferences). Never assume an English-tested pattern works on other scripts: run it on a few real rows of every language in your dataset and look at the output.

Another practical example is Arabic-script diacritics (the short-vowel marks in the Unicode range \u064B to \u065F). Removing them with flags=re.UNICODE makes spelling variants of the same Urdu or Arabic word match each other, which noticeably improves word counts and search.

Building a Reusable Cleaning Pipeline

Copying cleaning code from notebook to notebook leads to subtle differences between experiments. The professional approach is to define your cleaning steps as small, single-purpose functions and chain them with DataFrame.pipe(). The result reads like a recipe and is easy to test, reorder and reuse.

Small Functions, One Job Each

Python
import html

def drop_empty(df, col="text"):
mask = df[col].notna() & df[col].str.strip().ne("")
return df.loc[mask]

def drop_duplicate_text(df, col="text"):
return df.drop_duplicates(subset=col, keep="first")

def add_raw_features(df, col="text"):
return df.assign(
n_exclaim=df[col].str.count("!"),
n_caps_words=df[col].str.count(r"\b[A-Z]{2,}\b"),
has_url=df[col].str.contains(PATTERNS["url"], regex=True),
)

def clean_text(df, col="text", out="clean"):
s = (
df[col]
.map(html.unescape)
.str.normalize("NFKC")
.str.replace(PATTERNS["html"], " ", regex=True)
.str.replace(PATTERNS["url"], " <URL> ", regex=True)
.str.replace(PATTERNS["email"], " <EMAIL> ", regex=True)
.str.replace(PATTERNS["hashtag"], r"\1", regex=True)
.str.replace(PATTERNS["repeat_punct"], r"\1", regex=True)
.str.casefold()
.str.replace(PATTERNS["multi_space"], " ", regex=True)
.str.strip()
)
return df.assign(**{out: s})

Chaining with pipe()

Python
clean_df = (
df
.pipe(drop_empty)
.pipe(drop_duplicate_text)
.pipe(add_raw_features) # before lowercasing, to keep CAPS signal
.pipe(clean_text)
.reset_index(drop=True)
)
print(clean_df[["text", "clean"]].head())

Read the chain from top to bottom and you can explain the whole preprocessing strategy to a colleague in thirty seconds. That clarity is the real benefit of pipe().

Keep the original text

Always store the cleaned text in a new column and keep the raw column untouched. When a model behaves strangely, you will want to compare what the user actually wrote with what the model actually saw. Disk space is cheap; lost context is expensive.

Test Your Cleaning Functions

Cleaning code is real code and deserves tests. A handful of pytest cases protects you from regressions when you tweak a pattern months later:

Python
def test_clean_text_masks_urls():
df = pd.DataFrame({"text": ["Visit https://example.com NOW!!!"]})
out = clean_text(df)["clean"].iloc[0]
assert out == "visit <url> now!"

Notice the expected output contains <url> in lowercase, because casefold() runs after masking. Tests like this document your pipeline’s behaviour as clearly as any comment.

How Much Cleaning Is Too Much?

Here is an important nuance that many tutorials skip: modern transformer models need much less cleaning than classic models. Models like BERT, RoBERTa or today’s LLMs were trained on raw web text. They understand capitalisation, punctuation, emojis and even typos. Aggressive cleaning (lowercasing, removing punctuation, stemming) can actually hurt them by deleting information they know how to use.

Cleaning stepClassic ML (TF-IDF + linear model)Transformer / LLM
Remove HTML, boilerplate, encoding junkYesYes
Mask URLs, emails, personal dataYesYes (especially for privacy)
Normalise Unicode and whitespaceYesYes
LowercaseUsuallyNo
Remove punctuationOftenNo
Remove stopwordsOftenNo
Stemming / lemmatizationOftenNo

The rule of thumb: remove noise for everyone, remove information only for simple models. Keep two columns if you need both: a lightly cleaned text_model for transformers and a heavily normalised text_bow for bag-of-words methods.

Tokenization: Turning Sentences Into Pieces

Tokenization is the process of splitting text into units called tokens. Tokens are the atoms that every NLP method counts, compares or feeds into a model. Depending on the method, a token can be a word, a piece of a word, a punctuation mark or even a single character.

Pandas does not tokenize text by itself in any linguistic sense, but it gives you the perfect structure to store and analyse tokens. There are three levels of tokenization you should know.

Level 1: Whitespace and Regex Tokenization with Pandas

For quick analysis, splitting on whitespace or extracting word-like sequences with a regex is fast and often good enough:

Python
clean_df["tokens"] = clean_df["clean"].str.findall(r"<\w+>|\w+(?:'\w+)?")

This pattern keeps placeholder tokens like <url> intact, treats contractions like “don’t” as single tokens, and ignores punctuation. Each row of the tokens column now holds a Python list.

Level 2: Linguistic Tokenization with spaCy

Real language is messier than whitespace. “U.S.” is one token, “can’t” is arguably two, and “New York-based” is a puzzle. spaCy handles these cases with rules written by linguists, and it adds part-of-speech tags, lemmas and entities at the same time.

The most common mistake beginners make is calling spaCy row by row with apply(). That is slow because it bypasses spaCy’s batching. Instead, stream the whole column through nlp.pipe():

Python
import spacy

nlp = spacy.load("en_core_web_sm", disable=["parser", "ner"])

docs = nlp.pipe(clean_df["clean"], batch_size=500)
clean_df["lemmas"] = [
[tok.lemma_ for tok in doc if not tok.is_punct and not tok.is_space]
for doc in docs
]

Two optimisations are hiding in that snippet. First, disable=[...] switches off pipeline components you do not need, which can make spaCy several times faster. Second, nlp.pipe() processes texts in batches and can use multiple processes with the n_process argument for very large datasets.

Level 3: Subword Tokenization for Transformers

Transformer models use subword tokenizers (such as BPE, WordPiece or SentencePiece) that break rare words into frequent pieces. “Unbelievably” might become “un”, “believ”, “ably”. You do not choose these tokens; each model comes with its own tokenizer, and you must use the matching one.

Pandas is useful here for measuring token counts so you can choose a sensible maximum sequence length or estimate API costs:

Python
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("distilbert-base-uncased")
clean_df["n_tokens"] = [len(ids) for ids in tok(clean_df["text"].tolist())["input_ids"]]
print(clean_df["n_tokens"].describe())

Passing the whole list to the tokenizer at once is far faster than tokenizing row by row, because modern Hugging Face tokenizers are implemented in Rust and batch internally.

Counting tokens for LLM cost estimates

If you send text to a paid LLM API, cost is measured in tokens. A n_tokens column plus clean_df["n_tokens"].sum() gives you a reliable estimate before you spend money. Group by category to see which document types are most expensive.

Stopwords, Stemming and Lemmatization

Once text is tokenized, classic NLP pipelines usually simplify the vocabulary further. The goal is to make words that mean the same thing look the same, and to remove words that carry little meaning.

Removing Stopwords

Stopwords are extremely common words such as “the”, “is”, “and” or “of”. In a bag-of-words model they add noise and inflate the vocabulary. NLTK and spaCy both ship stopword lists:

Python
from nltk.corpus import stopwords

STOP = set(stopwords.words("english"))
KEEP = {"not", "no", "nor", "never", "very", "too"} # sentiment matters!
STOP -= KEEP

clean_df["tokens_nostop"] = clean_df["tokens"].map(
lambda toks: [t for t in toks if t not in STOP]
)

The KEEP set is the most important line in that block. Standard stopword lists include negations, and removing “not” turns “not good” into “good”, flipping the meaning of a review. Always review a stopword list against your task before trusting it.

Why use map() here?

The tokens column holds Python lists, and there is no vectorized .str method that filters list elements. A map() with a list comprehension is the clean, readable way to transform list-valued cells. For very large data, the explode-filter-regroup pattern shown later is often faster.

Stemming vs Lemmatization

Both techniques reduce words to a base form, but they work very differently.

StemmingLemmatization
How it worksChops word endings with rulesUses vocabulary and grammar
“studies”“studi”“study”
“better”“better”“good” (as adjective)
“running”“run”“run”
SpeedVery fastSlower
OutputMay not be a real wordAlways a real word
Typical toolNLTK PorterStemmer, SnowballStemmerspaCy, NLTK WordNet

Stemming is like trimming every word with the same pair of scissors: quick, but sometimes you cut too much. Lemmatization is like looking every word up in a dictionary: slower, but accurate. For search engines and very large corpora, stemming is often fine. For anything humans will read (topic labels, keyword reports), prefer lemmatization.

Python
from nltk.stem import SnowballStemmer

stemmer = SnowballStemmer("english")
clean_df["stems"] = clean_df["tokens_nostop"].map(
lambda toks: [stemmer.stem(t) for t in toks]
)

Word Frequencies and N-grams with explode()

This is where pandas truly shines for NLP. Once you have a column of token lists, a single method called explode() turns every list element into its own row. Suddenly, the entire toolkit of pandas aggregation (value_counts, groupby, pivot_table, merge) is available for words.

The Explode Pattern

Python
tokens_long = (
clean_df[["review_id", "label", "tokens_nostop"]]
.explode("tokens_nostop")
.rename(columns={"tokens_nostop": "token"})
.dropna(subset=["token"])
.reset_index(drop=True)
)
print(tokens_long.head())
# review_id label token
# 0 r1 positive absolutely
# 1 r1 positive love
# 2 r1 positive phone
# ...

By default explode() repeats the original row index for every token. The final reset_index(drop=True) gives each token row a unique index, which prevents confusing errors later in operations such as pd.crosstab().

Think of explode() as unfolding a stack of index cards: each review was one card holding many words; now each word is its own card, but it still remembers which review and which label it came from. This “long” or “tidy” format is the foundation for almost every text statistic.

Top Words Overall and Per Class

Python
top_overall = tokens_long["token"].value_counts().head(20)

top_by_label = (
tokens_long.groupby("label")["token"]
.value_counts()
.groupby(level=0)
.head(10)
)

Comparing the top words of positive and negative reviews is one of the fastest ways to understand why customers feel the way they do. In a real dataset, you might find “battery”, “fast” and “love” leading the positive side, and “support”, “broke” and “refund” leading the negative side.

Document Frequency vs Term Frequency

Raw counts can mislead you, because one long angry review that repeats “terrible” twenty times will dominate. Document frequency counts how many documents contain a word, regardless of how often:

Python
doc_freq = (
tokens_long.drop_duplicates(["review_id", "token"])
["token"].value_counts()
)

Document frequency is also the “DF” in TF-IDF, which we will use for machine learning in the next section.

Which Words Are Distinctive for Each Class?

A simple but powerful idea: compare how often each word appears in one class versus another. Words with a high ratio are distinctive.

Python
import numpy as np

counts = pd.crosstab(tokens_long["token"], tokens_long["label"])
rates = (counts + 1) / (counts.sum() + 1) # add-one smoothing
counts["log_ratio"] = np.log(rates["positive"] / rates["negative"])

print(counts.sort_values("log_ratio").head(10)) # most negative words
print(counts.sort_values("log_ratio").tail(10)) # most positive words

The + 1 smoothing prevents division by zero for words that appear in only one class. The log ratio is symmetrical: a word twice as common in positive reviews scores about +0.69, and twice as common in negative reviews scores about −0.69. This small table is often more insightful for stakeholders than any model accuracy number.

Building N-grams

Single words lose context. “Not good” and “customer service” mean more as pairs. N-grams are sequences of n consecutive tokens: bigrams (n = 2), trigrams (n = 3) and so on. A tiny helper with zip builds them from token lists:

Python
def ngrams(tokens, n=2):
return [" ".join(gram) for gram in zip(*(tokens[i:] for i in range(n)))]

clean_df["bigrams"] = clean_df["tokens"].map(ngrams)

top_bigrams = clean_df["bigrams"].explode().value_counts().head(15)

Notice that we built bigrams from tokens (with stopwords) rather than tokens_nostop. Phrases like “would buy again” or “not worth it” rely on small words, and removing stopwords first would destroy them.

Visualising Frequencies

A horizontal bar chart is the clearest way to present word frequencies:

Python
ax = top_overall.sort_values().plot.barh(figsize=(7, 6), title="Top 20 words")
ax.set_xlabel("Count")

Word clouds look attractive, but bar charts are easier to read accurately, because humans compare lengths far better than font sizes.

If your data has timestamps, combine exploded tokens with time grouping to see how topics rise and fall. For example, a spike in “delivery” complaints after a courier change is exactly the kind of insight businesses pay for:

Python
# assuming a 'created_at' datetime column
trend = (
tokens_long.merge(df[["review_id", "created_at"]], on="review_id")
.query("token in ['delivery', 'battery', 'refund']")
.groupby([pd.Grouper(key="created_at", freq="W"), "token"])
.size()
.unstack(fill_value=0)
)
trend.plot(title="Weekly mentions of key topics")

Feature Engineering: From Text to Numbers

Machine learning models do not read; they calculate. Feature engineering is the art of turning text into numbers that capture what matters for your task. Pandas is where those numbers live, side by side with labels and metadata, ready for modelling.

There are three families of text features, and strong models often combine all three.

Family 1: Handcrafted Statistical Features

These are features you design yourself from domain knowledge. They are cheap, interpretable and surprisingly strong, especially for tasks like spam detection, urgency prediction or quality scoring.

Python
def add_style_features(df, col="text"):
s = df[col]
n_chars = s.str.len()
n_words = s.str.split().str.len()
return df.assign(
n_chars=n_chars,
n_words=n_words,
avg_word_len=n_chars / n_words,
upper_ratio=s.str.count(r"[A-Z]") / n_chars,
punct_ratio=s.str.count(r"[^\w\s]") / n_chars,
n_digits=s.str.count(r"\d"),
starts_with_question=s.str.strip().str.match(r"(?i)(why|how|what|when|can|is)\b"),
)

Every one of these columns has a story. A high upper_ratio suggests shouting. A high punct_ratio may indicate excitement or spam. Many digits might mean order numbers, which hints at a support request. Because these features are just columns, you can immediately check their usefulness with a group comparison:

Python
features = add_style_features(clean_df)
print(features.groupby("label")[["upper_ratio", "n_words", "punct_ratio"]].mean())

If a feature’s average differs clearly between classes, it is likely to help a model.

Family 2: Bag-of-Words and TF-IDF

The classic way to represent a document is to count its words. TF-IDF (Term Frequency × Inverse Document Frequency) improves on raw counts by down-weighting words that appear in almost every document and up-weighting words that are distinctive.

An analogy helps. Imagine a library where every book contains the word “the”. Knowing that a book contains “the” tells you nothing. But if only three books contain “photosynthesis”, that word says a lot about those three books. TF-IDF is a mathematical version of that intuition.

Scikit-learn’s TfidfVectorizer takes a pandas column directly:

Python
from sklearn.feature_extraction.text import TfidfVectorizer

vec = TfidfVectorizer(
ngram_range=(1, 2), # words and bigrams
min_df=2, # ignore very rare terms
max_df=0.9, # ignore terms in more than 90% of documents
sublinear_tf=True, # dampen very frequent terms
)
X = vec.fit_transform(clean_df["clean"])
print(X.shape) # (n_documents, n_features), stored as a sparse matrix

The result is a sparse matrix: most entries are zero, because each document uses only a tiny fraction of the vocabulary. Sparse matrices save enormous amounts of memory. Avoid converting them into a dense DataFrame unless the vocabulary is small, but for inspecting the top terms of a single document, pandas is perfect:

Python
def top_terms(row_idx, n=8):
row = pd.Series(
X[row_idx].toarray().ravel(),
index=vec.get_feature_names_out(),
)
return row[row > 0].sort_values(ascending=False).head(n)

print(top_terms(0))

Family 3: Dense Embeddings

Embeddings represent each document as a dense vector of a few hundred numbers produced by a neural network. Unlike TF-IDF, embeddings capture meaning: “the phone dies quickly” and “battery life is poor” end up close together even though they share no words. We cover embeddings in detail in a dedicated section below.

End-to-End Project: A Sentiment Classifier

Time to put everything together. We will build a sentiment classifier on a realistic workflow: split the data, combine text and numeric features, train a model and evaluate it, all with pandas columns as the input.

Step 1: Split Without Leakage

Python
from sklearn.model_selection import train_test_split

train_df, test_df = train_test_split(
clean_df,
test_size=0.2,
stratify=clean_df["label"], # keep class balance in both sets
random_state=42,
)

Two habits matter here. stratify keeps label proportions equal in both splits. And because we removed duplicates before splitting, the same review cannot appear in both training and test sets, which would inflate test scores.

Fit on training data only

Vectorizers, scalers and stopword decisions learned from data must be fitted on the training set only. If you fit TF-IDF on all data and then split, information about the test set leaks into training. Scikit-learn pipelines prevent this automatically, which is why we use them next.

Step 2: Combine Text and Numeric Columns with ColumnTransformer

Scikit-learn’s ColumnTransformer lets one model consume several DataFrame columns, each with its own preprocessing. Here, the cleaned text goes through TF-IDF, while our handcrafted features are scaled:

Python
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

numeric_cols = ["n_exclaim", "n_caps_words"]

preprocess = ColumnTransformer([
("tfidf", TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True), "clean"),
("num", StandardScaler(), numeric_cols),
])

model = Pipeline([
("prep", preprocess),
("clf", LogisticRegression(max_iter=1000, class_weight="balanced")),
])

model.fit(train_df, train_df["label"])

Notice that we pass the DataFrame itself to fit(). The ColumnTransformer picks out the columns by name. This keeps your feature definitions next to your data and makes the pipeline easy to save and reuse with joblib.

Step 3: Evaluate Properly

Python
from sklearn.metrics import classification_report

preds = model.predict(test_df)
print(classification_report(test_df["label"], preds))

The report shows precision, recall and F1 for each class. On imbalanced data, look especially at the minority class: a model that never detects negative reviews is useless for a customer-service team, however high its overall accuracy.

Step 4: Error Analysis with Pandas

The most valuable step comes after the metrics. Put predictions back into a DataFrame and read the mistakes:

Python
results = test_df.assign(
pred=preds,
prob_positive=model.predict_proba(test_df)[:, list(model.classes_).index("positive")],
)
errors = results.loc[results["label"] != results["pred"], ["text", "label", "pred", "prob_positive"]]
print(errors.sort_values("prob_positive"))

Reading twenty misclassified examples will teach you more than any hyperparameter search. You will discover sarcasm (“Great, it broke on day one”), mixed reviews (“Camera is great but battery is awful”), and labelling errors in the data itself. Each discovery suggests a concrete next step: a new feature, a cleaning rule, a relabelling pass, or the need for a transformer model.

Step 5: Explain the Model

Linear models are wonderfully transparent. Pair each TF-IDF feature with its learned weight, and you can show stakeholders exactly which words push predictions in each direction:

Python
tfidf = model.named_steps["prep"].named_transformers_["tfidf"]
coefs = model.named_steps["clf"].coef_.ravel()[: len(tfidf.get_feature_names_out())]

weights = pd.Series(coefs, index=tfidf.get_feature_names_out()).sort_values()
print(weights.head(10)) # most negative weights push towards model.classes_[0] ("negative")
print(weights.tail(10)) # most positive weights push towards model.classes_[1] ("positive")

In a binary logistic regression, scikit-learn sorts class names alphabetically, so classes_[0] is “negative” and classes_[1] is “positive”. Positive weights pull a prediction towards the second class.

Enriching Rows: Sentiment Scores and Named Entities

Sometimes you do not need to train a model at all. Pre-built tools can add rich information to every row, and pandas makes the results instantly analysable.

Lexicon-Based Sentiment with VADER

VADER is a rule-based sentiment analyser tuned for social media. It understands capitalisation, exclamation marks, emoticons and negation, so run it on the raw text, not the cleaned version:

Python
from nltk.sentiment import SentimentIntensityAnalyzer

sia = SentimentIntensityAnalyzer()
scores = df["text"].fillna("").map(sia.polarity_scores).apply(pd.Series)
df = df.join(scores) # adds columns: neg, neu, pos, compound

The compound score ranges from −1 (very negative) to +1 (very positive). A common convention labels scores above 0.05 as positive and below −0.05 as negative. Lexicon methods will never match a fine-tuned transformer on accuracy, but they need no training data, run instantly and are perfect for a first look or a baseline.

Transformer Sentiment in One Line

When you need higher accuracy and have a GPU or a little patience, a Hugging Face pipeline can score an entire column. Pass a list, not single strings, so the model can batch:

Python
from transformers import pipeline

clf = pipeline("sentiment-analysis", model="distilbert-base-uncased-finetuned-sst-2-english")
outputs = clf(df["text"].fillna("").tolist(), batch_size=32, truncation=True)
df = df.join(pd.DataFrame(outputs).add_prefix("hf_")) # hf_label, hf_score

Named Entity Recognition (NER)

NER finds names of people, organisations, places, products, dates and money amounts. Storing entities in a long-format DataFrame makes them easy to count and filter:

Python
nlp_ner = spacy.load("en_core_web_sm")

rows = []
for review_id, doc in zip(df["review_id"], nlp_ner.pipe(df["text"].fillna(""))):
for ent in doc.ents:
rows.append({"review_id": review_id, "entity": ent.text, "type": ent.label_})

entities = pd.DataFrame(rows)
print(entities.groupby("type")["entity"].value_counts().head(20))

Building a list of dictionaries in a loop and creating the DataFrame once at the end is the efficient pattern. Appending to a DataFrame inside a loop is one of the slowest things you can do in pandas.

Embeddings and Semantic Search in a DataFrame

Embeddings are the bridge between classic data analysis and modern AI. Each document becomes a vector, and documents with similar meaning get similar vectors. With embeddings stored next to your text, pandas can power semantic search, clustering and duplicate detection.

Creating Embeddings

Python
from sentence_transformers import SentenceTransformer

encoder = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
emb = encoder.encode(
clean_df["text"].tolist(),
batch_size=64,
normalize_embeddings=True, # unit length, so dot product = cosine similarity
show_progress_bar=True,
)
print(emb.shape) # (n_documents, 384)

Keep the embedding matrix as a NumPy array rather than spreading 384 numbers across 384 DataFrame columns. Store a reference by row position, or save it alongside your data in Parquet as a list column if you need persistence.

Because the vectors are normalised, cosine similarity is a simple matrix product:

Python
def search(query, k=3):
q = encoder.encode([query], normalize_embeddings=True)[0]
sims = emb @ q
return (
clean_df.assign(score=sims)
.nlargest(k, "score")[["review_id", "text", "score"]]
)

print(search("the battery dies too quickly"))

This finds reviews about poor battery life even if they never use the word “dies”. For a few hundred thousand documents this brute-force approach is fast enough; beyond that, move the vectors into a vector index such as FAISS or a vector database.

Finding Semantic Duplicates

Earlier we found exact and normalised duplicates. Embeddings catch paraphrases too:

Python
sim = emb @ emb.T
i, j = np.triu_indices_from(sim, k=1)
pairs = pd.DataFrame({"a": i, "b": j, "sim": sim[i, j]})
near_dupes = pairs.query("sim > 0.92").sort_values("sim", ascending=False)

The full similarity matrix grows with the square of the number of documents, so this exact approach suits datasets up to a few tens of thousands of rows. For larger data, use approximate nearest-neighbour search.

Performance: Processing Millions of Rows

A pipeline that takes two seconds on 1,000 rows can take hours on 10 million. The good news is that a few habits deliver most of the speed.

1. Prefer Vectorized .str Methods Over apply()

Python
# Slow: a Python function call for every row
df["n_words"] = df["text"].apply(lambda t: len(t.split()))

# Fast: vectorized, and Arrow-accelerated with the str dtype in pandas 3.x
df["n_words"] = df["text"].str.split().str.len()

With pandas 3.x and PyArrow installed, many .str methods execute inside Arrow’s compiled code. Benchmarks published around the 3.0 release reported large speed-ups for common string operations compared with the old object dtype.

2. Combine Patterns Into One Pass

Running ten separate str.replace() calls scans the column ten times. When replacements share the same target, combine them with |:

Python
NOISE = "|".join([PATTERNS["html"], PATTERNS["mention"]])
df["clean"] = df["text"].str.replace(NOISE, " ", regex=True)

3. Use Categories for Repeated Labels

Columns like label, lang or source have few distinct values repeated millions of times. Converting them to the category dtype can shrink memory dramatically and speed up groupby:

Python
df["label"] = df["label"].astype("category")
print(df.memory_usage(deep=True).sum() / 1e6, "MB")

4. Batch Your NLP Libraries

spaCy’s nlp.pipe(), Hugging Face tokenizers, transformer pipelines and sentence-transformers all accept lists and process them in batches. Passing whole columns as lists instead of calling the model per row is frequently the single biggest speed improvement in an NLP pipeline.

5. Process in Chunks and Save to Parquet

For files larger than memory, process chunk by chunk and append results to partitioned Parquet files:

Python
for i, chunk in enumerate(pd.read_csv("huge.csv", chunksize=200_000)):
out = chunk.pipe(drop_empty).pipe(clean_text)
out.to_parquet(f"clean/part-{i:04d}.parquet", index=False)

full = pd.read_parquet("clean/") # reads every part as one DataFrame

6. Know When to Reach for Another Tool

Pandas is excellent up to roughly the size of your RAM. Beyond that, or when you need multi-core speed for every operation, consider Polars (a fast DataFrame library with a lazy query engine), DuckDB (SQL over Parquet files) or Hugging Face Datasets (memory-mapped Arrow tables with a map() method designed for NLP). All three speak Arrow, so with pandas 3.x moving data between them is cheaper than ever.

Data sizeRecommended approach
Up to a few hundred MBPlain pandas, vectorized .str methods
Close to or above RAMpandas with chunks + Parquet, or DuckDB
Many GB, heavy transformationsPolars (lazy mode) or DuckDB
Model training at scaleHugging Face Datasets with batched map()

Best Practices and Common Mistakes

After years of text projects, the same mistakes appear again and again. Here is a checklist to keep you on the right side of them.

Do This

  • Profile before you clean. Run your profiling function on every new dataset.
  • Keep raw text. Store each transformation in a new column.
  • Write cleaning as small functions chained with pipe(), and test them.
  • Deduplicate before splitting into training and test sets.
  • Match cleaning to the model. Light cleaning for transformers, heavier normalisation for bag-of-words models.
  • Use na=False in str.contains() when building filter masks.
  • Always pass regex=True to str.replace() when using patterns.
  • Batch every model call with lists or nlp.pipe().
  • Read your errors. Error analysis beats blind hyperparameter tuning.

Avoid This

  • Removing negations with a default stopword list in sentiment tasks.
  • Lowercasing before extracting case-based features.
  • Chained assignment like df["text"][mask] = ..., which silently fails under Copy-on-Write.
  • Appending to DataFrames in a loop. Build a list, then create the DataFrame once.
  • Checking dtype == object to find text columns in pandas 3.x.
  • Exploding huge columns blindly. One million documents of 200 tokens become 200 million rows; filter or sample first.
  • Storing personal data in clear text. Mask emails, phone numbers and names before sharing datasets.

Pandas NLP Cheat Sheet

Bookmark this table. It covers the operations you will use in almost every project.

TaskPandas code
Count missing textdf["text"].isna().sum()
Count blank textdf["text"].str.strip().eq("").sum()
Word countdf["text"].str.split().str.len()
Lowercase (multilingual)df["text"].str.casefold()
Normalise Unicodedf["text"].str.normalize("NFKC")
Remove URLsdf["text"].str.replace(r"https?://\S+", "", regex=True)
Collapse whitespacedf["text"].str.replace(r"\s+", " ", regex=True).str.strip()
Filter by keyworddf[df["text"].str.contains("refund", case=False, na=False)]
Extract all hashtagsdf["text"].str.findall(r"#(\w+)")
Tokens to rowsdf.explode("tokens")
Top wordsdf["tokens"].explode().value_counts().head(20)
Words per classlong.groupby("label")["token"].value_counts()
One-hot tagsdf["tags"].str.get_dummies(sep="|")
Drop duplicatesdf.drop_duplicates(subset="text")
New columns cleanlydf.assign(n=pd.col("text").str.len())
Save processed datadf.to_parquet("clean.parquet", index=False)

Frequently Asked Questions

Is pandas good for NLP?

Yes. Pandas is not an NLP library in itself, but it is the best tool in Python for organising, cleaning, exploring and analysing text datasets. It integrates directly with scikit-learn, spaCy, NLTK and Hugging Face, which makes it the natural backbone of most NLP workflows.

Should I use pandas or Polars for text processing in 2026?

Use pandas when your data fits comfortably in memory, when you rely on its huge ecosystem, or when you are following tutorials and team conventions. Choose Polars when you need maximum speed on large datasets or lazy evaluation. Since pandas 3.x stores strings in Arrow format, many workflows can combine both with little conversion cost.

Do I need to clean text for large language models?

Much less than for classic models. Remove genuine noise (HTML, encoding errors, boilerplate, duplicates) and mask personal data, but keep capitalisation, punctuation and stopwords. LLMs were trained on natural text and use those signals.

Why does str.replace() not remove my pattern?

Since pandas 2.0, str.replace() treats the pattern as a literal string unless you pass regex=True. Add that argument whenever your pattern contains regex syntax such as \d, + or |.

How do I apply spaCy to a pandas column efficiently?

Use nlp.pipe(df["text"]) instead of df["text"].apply(nlp). Disable unused components with spacy.load(..., disable=[...]), increase batch_size, and use n_process for multi-core processing on large datasets.

What is the difference between object and str dtype in pandas?

object can store any Python object, so it is flexible but slow and ambiguous. The str dtype, the default for text in pandas 3.x, stores only strings and missing values, uses Arrow under the hood when PyArrow is installed, and makes string operations faster and more memory-efficient.

Conclusion: Your Text Workbench Is Ready

You have now walked through a complete, modern NLP data workflow built on pandas: loading text from any format, profiling it in seconds, cleaning it with vectorized string methods and regex, normalising Unicode and handling multiple languages, tokenizing and lemmatizing at scale, counting words and n-grams with explode(), engineering features, training and explaining a classifier, enriching rows with sentiment, entities and embeddings, and keeping everything fast on millions of rows.

The deeper lesson is about mindset. Models get the headlines, but data quality decides results. The developer who can confidently inspect, clean and reshape text, and who knows when not to clean it, will build better NLP systems than someone who only knows how to call a model. Pandas gives you that confidence.

Here is a practical next step: take a text dataset you care about (product reviews, support tickets, social media comments in your own language) and run the full pipeline from this guide on it. Profile it, clean it, explode it, model it and, most importantly, read the errors. One real project will teach you more than ten tutorials.

Information reflects the state of tools, libraries and plans as of October 2026. AI products change quickly, so check official documentation and pricing pages for the latest details.

Comments

Popular posts from this blog

PyTorch Explained: The Complete Guide to Deep Learning & Neural Networks in 2026

  PyTorch: The Complete Guide to Deep Learning's Most Popular Framework Introduction If you've trained a neural network, fine-tuned a language model, or experimented with a diffusion-based image generator in the last several years, there's a strong chance PyTorch was somewhere underneath it. Originally released by Facebook AI Research (now Meta AI) in 2016, PyTorch has grown from a research-focused alternative to established frameworks into the dominant tool in the deep learning world — powering everything from academic papers to some of the largest AI systems ever deployed in production. This guide takes a deep, practical look at PyTorch: what it is, why it was designed the way it was, how its core components fit together, and how to actually use it to build, train, and deploy real models. Whether you're completely new to deep learning or you've used other frameworks and want to understand what makes PyTorch different, this article will walk you through everythi...

Multimodal AI Explained: How Text, Image & Voice Merge Into One Model

  The Rise of Multimodal AI: Text, Image, and Voice in One Model For most of the last decade, AI systems were narrow specialists: a language model that only understood text, an image classifier that only understood pictures, a speech recognition system that only understood audio. Getting these systems to work together meant stitching together separate pipelines, converting between formats at every handoff, and accepting the errors and awkwardness that came with each conversion step. That era is ending. Modern multimodal AI systems process text, images, audio, and increasingly video within a single unified model, reasoning across all of them together rather than treating each as a separate problem solved by a separate system. This guide explains what multimodal AI actually is, the real architectural difference between "bolted-together" and "natively unified" multimodal systems, where this technology is already changing real products, and an honest look at where it ...

How RAG Works in AI: Retrieval-Augmented Generation Explained

How RAG Works in AI: Retrieval-Augmented Generation Explained Introduction Ask a large language model about something that happened last week, or about the contents of your company's internal wiki, or about a product manual that was never part of its training data, and you'll run into the same wall every time: the model simply doesn't know. It wasn't trained on that information, and no amount of clever prompting can make it recall a fact it never saw. Retrieval-Augmented Generation, almost universally shortened to RAG, is the technique that solves this problem, and it has quietly become one of the most widely deployed patterns in production AI systems — powering everything from customer support chatbots that answer questions using a company's own documentation, to coding assistants that search a codebase before answering, to research tools that cite specific passages from specific documents rather than answering from memory alone. This article explains what RAG actu...