NLP Sentiment Analysis: A Practical Guide from Lexicons to LLMs
· @Syed Wahab Uddin
Introduction: What Sentiment Analysis Is and Why It Matters
Sentiment analysis is the NLP task of identifying the opinion, attitude or emotion expressed in text. At its simplest, it answers one question: is this text positive, negative or neutral? Also called opinion mining, it turns huge volumes of unstructured reviews, posts and messages into numbers a team can act on.
Consider three everyday examples:
- "Delivery was quick and the packaging was perfect." is positive.
- "The app crashes every time I open my cart." is negative.
- "The order arrived on Tuesday." is neutral.
A person labels these in a second. Doing it reliably for 50,000 reviews a day, in several languages, full of slang and sarcasm, is where NLP comes in.
Why organizations invest in it
Most of what customers think about a product is written down somewhere: app store reviews, support tickets, survey comments, social posts and chat logs. Very little of it is ever read. Sentiment analysis makes that feedback measurable, so teams can track it, compare it and act on it.
Typical uses include:
- Brand monitoring: catching a spike in negative posts before it turns into a crisis
- Product feedback: learning which features customers love and which frustrate them
- Customer support: routing angry messages to senior agents first
- Market research: comparing how people talk about you and your competitors
- Finance: gauging market mood from news and social media
- Public services: understanding reactions to announcements and policies
A short history
Early systems in the 2000s relied on word lists: count the positive words, subtract the negative ones. Machine learning followed, with classifiers such as Naive Bayes and support vector machines trained on labeled reviews; a widely cited 2002 study by Pang, Lee and Vaithyanathan applied them to movie reviews. Deep learning added models that read word order, and transformers such as BERT pushed accuracy further from 2018 onward.
Today, large language models can classify sentiment from a plain-English instruction with no training data at all. Yet every earlier approach is still in use, because each fits different limits on budget, speed, accuracy and explainability. A word list may be enough for a quick dashboard, while a fine-tuned transformer may be worth it for a product making thousands of decisions per hour.
What this guide covers
This guide goes from concepts to working Python code. You will learn the main types of sentiment analysis, how to collect and label data, and how to build models with lexicons, classical machine learning, deep learning, transformers and LLMs. It closes with evaluation, the hardest open problems, and how to deploy a sentiment system that keeps working after launch.
Types of Sentiment Analysis
"Sentiment analysis" covers several related tasks. Picking the right one up front saves a lot of rework, because each needs different labels, data and models.
Type | What it outputs | Example | Best for |
|---|---|---|---|
Polarity | Positive, negative or neutral | "Loved it, will buy again" → positive | Review scoring, brand dashboards |
Fine-grained | A scale, such as 1 to 5 stars | "Decent, but overpriced" → 3 of 5 | Rating prediction, trend tracking |
Emotion detection | Joy, anger, sadness, fear, surprise and similar | "They cancelled my order again!" → anger | Support triage, content moderation |
Aspect-based | Sentiment per feature or target | "Great camera, awful battery" → camera positive, battery negative | Product teams, competitor analysis |
Intent detection | What the writer plans or wants | "Thinking of switching to another bank" → churn risk | Retention, sales follow-up |
Subjectivity detection | Opinion or fact | "The phone has 128 GB" → fact | Filtering text before sentiment scoring |
Document, sentence or aspect level
Granularity matters as much as the label set. Document-level analysis gives a whole text one label, sentence-level labels each sentence, and aspect-level labels opinions about specific targets. A hotel review that praises the location but complains about the room is "mixed" overall; aspect-level analysis tells the hotel exactly what to fix.
The tricky neutral class
Neutral is the hardest label to define. It gets mixed up with factual statements ("arrived Tuesday"), mixed opinions ("good food, slow service") and text that is simply unclear. Decide early whether "mixed" is its own class, because vague definitions are a top cause of annotators disagreeing with each other.
Scores versus labels
Many systems output a continuous score, such as −1 to +1, instead of or alongside a label. Scores let you set different thresholds for different uses and average them over time into trend lines. When decisions depend on confidence, such as auto-escalating a ticket, calibrated probabilities are worth more than raw scores.
Which type do you need?
Start from the decision the output will feed. A manager who needs a weekly trend line only needs polarity. A product team deciding what to fix needs aspect-based analysis, and a support team setting priorities is better served by emotion or urgency detection than by plain positive and negative.
How Sentiment Analysis Works
Every sentiment system, from a word list to a large language model, follows the same pipeline: gather text, prepare it, score it and turn the scores into decisions. The method in the middle changes; the steps around it do not.
Here is what happens at each stage:
- Text sources. Pull text from wherever customers speak: review sites, app stores, support tools, social platforms and surveys. Keep metadata such as date, product, channel and language, because you will want to slice results by them later.
- Preprocessing. Remove noise like HTML and duplicates, but keep the cues that carry feeling. A later section explains why this differs from general-purpose cleaning.
- Scoring. Pick one of four method families. Lexicons need no training, classical models learn from labeled data, transformers bring pretrained language understanding, and LLMs follow written instructions.
- Output. Each text gets a label such as positive, negative or neutral, a confidence score, and optionally aspects and emotions.
- Decide and act. Outputs feed dashboards, trigger alerts on spikes, route support tickets and inform product roadmaps.
- Review and retrain. People check a sample of predictions, especially low-confidence ones. Their corrections become new training data, so the model keeps up with new products, slang and topics.
Many production systems combine methods. A common pattern runs a fast, cheap model on everything and sends only low-confidence or high-value texts to an LLM or a human reviewer. The rest of this guide walks through each stage, starting with the data.
Data: Collecting and Labeling Sentiment Datasets
A sentiment model is only as good as its labeled data. Many failed projects trace back to training text that does not match what the model sees in production, or to labels that annotators never agreed on.
Start with public datasets
Public benchmarks are great for learning, prototyping and comparing methods. Sizes below are approximate.
Dataset | Text | Labels | Size |
|---|---|---|---|
IMDb Large Movie Review | Movie reviews | Positive or negative | About 50,000 reviews |
Stanford Sentiment Treebank (SST-2, SST-5) | Movie review sentences | 2 or 5 classes, with phrase-level labels | About 11,800 sentences |
Sentiment140 | Tweets | Positive or negative, labeled from emoticons | About 1.6 million tweets |
GoEmotions | Reddit comments | 27 emotions plus neutral | About 58,000 comments |
Amazon and Yelp reviews | Product and business reviews | 1 to 5 stars | Millions of reviews |
SemEval shared tasks | Tweets, restaurant and laptop reviews | 3 classes and aspect labels | Thousands per task |
Collect data from your own domain
Public data rarely matches your domain. A model trained on movie reviews learns that "unpredictable" is praise for a plot, yet "unpredictable steering" in a car review is a serious complaint. Collect real text from your own reviews, tickets and surveys, follow each platform's terms and privacy law, and strip personal information before anyone labels it.
Weak labels from ratings and emoticons
Star ratings give you labels for free. Map 1 to 2 stars to negative and 4 to 5 to positive, then treat 3 stars as neutral or drop it. These labels are noisy, since some people leave five stars with a complaint, but large volumes of noisy labels often train well, as long as your test set is labeled by hand.
Write annotation guidelines first
Before anyone labels by hand, write down the rules:
- Define each class with examples, including edge cases: mixed reviews, sarcasm, questions and complaints about a competitor.
- State whose sentiment counts. It is the writer's, so "the villain was terrifying" is positive in a horror film review.
- Decide how to handle spam, off-topic text and other languages.
Then run a pilot. Have two or three people label the same 100 to 200 examples, and measure agreement with Cohen's kappa for two annotators or Krippendorff's alpha for more. Low agreement, roughly below 0.6 kappa, usually means the guidelines need fixing, not the annotators.
How much data do you need?
Fine-tuning a pretrained transformer often gives strong results with a few thousand labeled examples per class, and careful labeling can make a few hundred work. Classical models usually need more. Coverage of hard cases matters more than raw volume, and active learning, which labels the examples the current model is least sure about, stretches a labeling budget further.
Imbalance and splits
Real data is usually imbalanced; on many review sites, positive reviews dominate. Keep the natural class mix in your test set so metrics reflect reality, and handle imbalance during training with class weights or resampling. Where possible, split by time, product or user, so near-duplicate texts cannot leak from training into testing.
Preprocessing Text for Sentiment
General-purpose preprocessing recipes often hurt sentiment models, because they delete the very words and symbols that carry feeling. The rule for sentiment is simple: clean the noise, keep the emotion.
Keep negations
Negation flips meaning: "good" versus "not good," or "I would recommend" versus "I would never recommend." Default stop-word lists in NLTK and spaCy include "not," "no" and "never," so take those words off the list or skip stop-word removal entirely.
For bag-of-words models, a classic trick marks the scope of a negation. Every word after a negation gets a _NEG suffix until the next punctuation mark, so "I did not like the ending." becomes "I did not like_NEG the_NEG ending_NEG." The model then learns "like_NEG" as a negative feature, and NLTK provides this as mark_negation.
Keep emojis, emoticons and punctuation
"Great :)" and "Great :(" mean opposite things, and an emoji often carries the whole message. Convert emojis to words with emoji.demojize, map emoticons like ":)" and ":-(" to tokens, and keep "!" and "?". Repeated exclamation marks are a strong signal of intensity.
Respect intensifiers and contrast
Words like "very," "extremely" and "absolutely" strengthen sentiment, while "slightly," "somewhat" and "a bit" soften it. Contrast words such as "but" and "however" shift weight to the clause that follows, so "The screen is beautiful, but the battery is terrible" leans negative. Lexicon tools apply these rules directly, and machine learning models learn them, provided the words are kept.
Case and elongation
ALL CAPS often signals strong emotion, as in "THIS IS THE WORST." If you lowercase for a classical model, consider adding a feature for the share of capitalized words. Squash elongations so "sooooo good" becomes "soo good," letting variants share one feature while keeping the emphasis.
Domain slang and mixed languages
Sentiment slang changes fast and varies by community. "Sick," "fire" and "insane" can all be compliments, while "mid" is a mild insult. In code-mixed text such as Roman Urdu, "bohat acha product hai" ("it's a very good product") contains no English sentiment word at all, so build a small domain dictionary from your own data.
For transformers, do less
Fine-tuned transformers learn negation, emojis and intensity from context. Give them close-to-raw text: remove HTML, URLs and duplicates, and leave everything else alone.
A preprocessing function for classical models
import re
import emoji
from nltk.sentiment.util import mark_negation
EMOTICONS = {":)": " smile_emo ", ":-)": " smile_emo ",
":(": " sad_emo ", ":-(": " sad_emo "}
def prep_for_sentiment(text: str) -> list[str]:
text = re.sub(r"https?://\S+", " URL ", text)
for emo, token in EMOTICONS.items():
text = text.replace(emo, token)
text = emoji.demojize(text, delimiters=(" ", " "))
text = re.sub(r"([a-zA-Z])\1{2,}", r"\1\1", text) # sooooo -> soo
tokens = re.findall(r"[\w']+|[!?.,]", text.lower()) # keeps "didn't" whole
return mark_negation(tokens)
print(prep_for_sentiment("I didn't like the ending. Sooooo slow :("))Lexicon-Based Methods: VADER, TextBlob and SentiWordNet
Lexicon-based sentiment analysis needs no training data. It looks words up in a dictionary of sentiment scores and combines them with hand-written rules. It is the fastest way to a first result, and it remains useful when you have no labeled data at all.
How lexicons work
A sentiment lexicon gives each word a score, so "excellent" might be strongly positive and "bad" moderately negative. The system adds up the scores of the words in a text, then adjusts them with rules for negation, intensifiers, punctuation and contrast. The final score can be thresholded into labels.
VADER
VADER (Valence Aware Dictionary and sEntiment Reasoner), introduced by Hutto and Gilbert in 2014, was built for social media. Its lexicon covers slang, emoticons and acronyms like "lol," and its rules handle capitals, exclamation marks, degree words such as "extremely," "but" contrasts and negation. It returns positive, negative and neutral proportions, plus a normalized compound score from −1 to +1.
A common convention labels a compound score of 0.05 or more as positive, −0.05 or less as negative, and anything between as neutral.
import nltk
from nltk.sentiment import SentimentIntensityAnalyzer
nltk.download("vader_lexicon")
sia = SentimentIntensityAnalyzer()
for text in ["The food was good.",
"The food was GOOD!!!",
"The food was not good.",
"The food was good, but the service was awful."]:
print(f"{sia.polarity_scores(text)['compound']:+.3f} {text}")Run it and the score climbs with capitals and exclamation marks, flips for "not good," and turns negative for the "but" sentence, because VADER gives more weight to the clause after "but."
TextBlob
TextBlob returns a polarity score from −1 to +1 and a subjectivity score from 0 (objective) to 1 (subjective). Its default analyzer uses an adjective lexicon from the Pattern library, built from product reviews. It is simple and good for teaching, but it misses slang and does worse than VADER on social text.
from textblob import TextBlob
blob = TextBlob("The plot was predictable but the acting was brilliant.")
print(blob.sentiment) # Sentiment(polarity=..., subjectivity=...)SentiWordNet
SentiWordNet scores the positivity, negativity and objectivity of WordNet synsets, meaning individual senses of words rather than words. That is more precise in theory, but you must first work out which sense is meant, which is a hard problem in its own right. It is mostly used in research and as extra features for other models.
Strengths and weaknesses
Lexicons need no labels or training, run fast and cheaply anywhere, and are fully explainable: you can show exactly which words drove a score. Their weaknesses are just as clear:
- Domain blindness. "Unpredictable" gets the same score for a thriller and a car.
- No sense of context. They miss sarcasm and implicit sentiment, such as "the battery lasted two hours."
- Language coverage. Most strong lexicons are English; others are smaller or missing.
- Accuracy. They usually trail trained models clearly on in-domain test sets.
Use lexicons for quick exploration, as a baseline, as features in a bigger model, or when you truly have no labels. You can also teach them your domain, since VADER's lexicon is a plain dictionary: sia.lexicon.update({"mid": -1.0, "fire": 2.0}).
Classical Machine Learning: Naive Bayes, Logistic Regression and SVMs
Supervised machine learning learns which features signal sentiment from labeled examples, instead of relying on a fixed dictionary. With TF-IDF features, classical classifiers train in seconds, adapt to your domain and make strong baselines.
The recipe
- Preprocess the text, keeping negations and emojis.
- Turn each document into a TF-IDF vector of word unigrams and bigrams, so "not good" becomes a feature of its own.
- Train a linear classifier.
- Evaluate on a held-out test set and read the errors.
Choosing a classifier
- Multinomial Naive Bayes assumes words are independent and estimates how likely each word is under each class. It is very fast and works well on small datasets and short texts, though its probabilities tend to be overconfident.
- Logistic regression learns one weight per feature and outputs well-behaved probabilities. Its weights are easy to read, so you can list the words that push hardest in each direction.
- Linear SVM finds the boundary with the widest margin between classes. It is often the most accurate of the three on text, but it needs calibration to output probabilities.
A well-known 2012 study by Wang and Manning showed that these simple linear models with bigrams, plus a hybrid of Naive Bayes and SVM, were hard to beat on the sentiment benchmarks of that time.
A complete classical model
from datasets import load_dataset
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report
from sklearn.pipeline import make_pipeline
imdb = load_dataset("imdb")
train, test = imdb["train"], imdb["test"]
model = make_pipeline(
TfidfVectorizer(ngram_range=(1, 2), min_df=3, max_df=0.9, sublinear_tf=True),
LogisticRegression(C=4.0, max_iter=2000),
)
model.fit(train["text"], train["label"])
print(classification_report(test["label"], model.predict(test["text"]),
target_names=["negative", "positive"]))On IMDb, a setup like this commonly reaches roughly 90% accuracy, with no GPU and a few minutes of work.
Look inside the model
import numpy as np
vec = model.named_steps["tfidfvectorizer"]
clf = model.named_steps["logisticregression"]
names = vec.get_feature_names_out()
order = np.argsort(clf.coef_[0])
print("Most negative:", names[order[:15]])
print("Most positive:", names[order[-15:]])Expect words like "worst," "waste" and "awful" on one side and "excellent," "perfect" and "great" on the other. This list doubles as a sanity check: if a product name or a username appears near the top, your data contains a shortcut the model has learned to exploit.
Tuning tips
- Search over
C, the n-gram range andmin_dfwithGridSearchCV. - For noisy or misspelled text, try character n-grams (
analyzer="char_wb", ngram_range=(2, 5)), or combine word and character features withFeatureUnion. - Set
class_weight="balanced"when classes are skewed. - Wrap an SVM in
CalibratedClassifierCVwhen you need probabilities.
When classical is enough
If your domain has clear sentiment vocabulary, you need explanations, or you must process huge volumes on cheap CPUs, a classical model may be all you need. Its weak spots are word order beyond bigrams, negation that spans a long clause, sarcasm and vocabulary it never saw in training.
Deep Learning: CNNs, RNNs and LSTMs
Before transformers, neural networks brought two big gains to sentiment analysis: dense word embeddings that capture similarity between words, and architectures that read word order.
Word embeddings as input
Instead of sparse counts, each word becomes a dense vector, often initialized from pretrained Word2Vec, GloVe or fastText embeddings. "Terrible" and "awful" start close together, so a model that learns from one generalizes to the other, even if only one appeared in training.
Convolutional neural networks
Text CNNs slide filters over windows of two to five words, acting as learned detectors for phrases like "not worth" or "highly recommend." Max pooling keeps the strongest signal from each filter. Yoon Kim's 2014 paper showed that a simple one-layer CNN over pretrained word vectors performed strongly across several sentence classification benchmarks.
CNNs are fast and good at local patterns. They struggle when the words that decide sentiment sit far apart.
Recurrent networks and LSTMs
Recurrent neural networks (RNNs) read text one token at a time, carrying a hidden state that summarizes everything so far. Plain RNNs forget quickly, so Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) cells add gates that decide what to keep and what to drop. Bidirectional LSTMs read in both directions, which helps with sentences like "The movie, which everyone said was amazing, was a letdown."
An attention layer on top of an LSTM lets the model weight the most important words. As a bonus, those weights give a rough view of which words drove each prediction.
A minimal Keras model
import numpy as np
import tensorflow as tf
from tensorflow.keras import layers
texts = np.array(train_texts) # review strings
labels = np.array(train_labels) # 0 = negative, 1 = positive
vectorize = layers.TextVectorization(max_tokens=30000, output_sequence_length=300)
vectorize.adapt(texts)
model = tf.keras.Sequential([
tf.keras.Input(shape=(), dtype="string"),
vectorize,
layers.Embedding(30000, 128, mask_zero=True),
layers.Bidirectional(layers.LSTM(64)),
layers.Dropout(0.3),
layers.Dense(1, activation="sigmoid"),
])
model.compile(optimizer="adam", loss="binary_crossentropy", metrics=["accuracy"])
model.fit(texts, labels, validation_split=0.1, epochs=3, batch_size=64)Where they stand today
CNNs and LSTMs train from scratch on modest hardware and still suit mobile and embedded use. For most new projects, though, fine-tuning a pretrained transformer gives higher accuracy with less labeled data, because the model already understands language before it sees your first example.
Transformers: Fine-Tuning BERT and RoBERTa
Transformers are the default choice for accurate sentiment analysis today. Models such as BERT and RoBERTa are pretrained on huge amounts of text, so they already handle grammar, negation and context. Fine-tuning adapts that knowledge to your labels with relatively little data.
Why transformers work so well
Self-attention lets every word look at every other word in the text at once. In "I thought I would hate it, but honestly it was wonderful," the model can tie "wonderful" to the overall verdict and see that "hate" describes an expectation that never came true. Contextual embeddings also handle words whose sentiment depends on context, like "cheap," which is good for a price and bad for build quality.
Start with a ready-made model
Before training anything, try a model already fine-tuned for sentiment. The Hugging Face Hub hosts many, including:
distilbert-base-uncased-finetuned-sst-2-english: fast positive-or-negative classification for Englishcardiffnlp/twitter-roberta-base-sentiment-latest: negative, neutral or positive, trained on tweetsnlptown/bert-base-multilingual-uncased-sentiment: 1 to 5 stars for product reviews in several European languagesj-hartmann/emotion-english-distilroberta-base: Ekman's six basic emotions plus neutral
from transformers import pipeline
clf = pipeline("sentiment-analysis",
model="cardiffnlp/twitter-roberta-base-sentiment-latest")
print(clf(["Customer service fixed it in five minutes, impressed!",
"Third time my order is late. Done with this app."]))If a ready-made model scores well on a sample of your own labeled data, you may not need to train at all.
Fine-tuning on your own data
When off-the-shelf accuracy is not enough, fine-tune. This example uses the Hugging Face Trainer and keeps the real test set untouched for the final score:
import numpy as np
from datasets import load_dataset
from sklearn.metrics import accuracy_score, f1_score
from transformers import (AutoModelForSequenceClassification, AutoTokenizer,
DataCollatorWithPadding, Trainer, TrainingArguments)
model_name = "distilroberta-base"
dataset = load_dataset("imdb")
tokenizer = AutoTokenizer.from_pretrained(model_name)
def tokenize(batch):
return tokenizer(batch["text"], truncation=True, max_length=256)
tokenized = dataset.map(tokenize, batched=True)
# A 6,000-review subset keeps training short; use everything for best results
splits = tokenized["train"].shuffle(seed=42).select(range(6000)).train_test_split(
test_size=0.15, seed=42)
model = AutoModelForSequenceClassification.from_pretrained(
model_name, num_labels=2,
id2label={0: "negative", 1: "positive"}, label2id={"negative": 0, "positive": 1})
def compute_metrics(eval_pred):
logits, labels = eval_pred
preds = np.argmax(logits, axis=-1)
return {"accuracy": accuracy_score(labels, preds),
"macro_f1": f1_score(labels, preds, average="macro")}
args = TrainingArguments(
output_dir="sentiment-model",
learning_rate=2e-5,
per_device_train_batch_size=16,
per_device_eval_batch_size=32,
num_train_epochs=2,
weight_decay=0.01,
eval_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
metric_for_best_model="macro_f1",
)
trainer = Trainer(model=model, args=args,
train_dataset=splits["train"], eval_dataset=splits["test"],
data_collator=DataCollatorWithPadding(tokenizer),
compute_metrics=compute_metrics)
trainer.train()
test_sample = tokenized["test"].shuffle(seed=42).select(range(5000)) # IMDb is sorted by label
print(trainer.evaluate(test_sample)) # held-out test
trainer.save_model("sentiment-model")
tokenizer.save_pretrained("sentiment-model")Practical tips
- Learning rate. Stay roughly between 1e-5 and 5e-5. Set it too high and the model forgets what it learned in pretraining.
- Epochs. Two to four is usually enough; watch validation scores for overfitting.
- Length. Most reviews fit in 256 tokens. For longer texts, keep the beginning and end, or split into chunks and average the predictions.
- Model size. DistilBERT and DistilRoBERTa are smaller and faster, while RoBERTa-base and DeBERTa-v3 are generally more accurate. Let your latency budget decide.
- Other languages. XLM-RoBERTa covers about 100 languages and fine-tunes the same way.
- Imbalance. Use class weights in a custom loss, or oversample the minority class.
Hardware and cost
A free cloud GPU, such as the ones Google Colab offers, can typically fine-tune a base-size model on a few thousand examples in well under an hour. For high-volume inference, export the model to ONNX or quantize it to cut latency and cost.
Large Language Models: Zero-Shot and Few-Shot Sentiment
Large language models can classify sentiment from an instruction alone. You describe the task in plain language, optionally add a few examples, and the model returns a label, with no training run and no training set.
Zero-shot and few-shot
- Zero-shot prompts contain only the instructions and the text to classify.
- Few-shot prompts add a handful of labeled examples, ideally covering edge cases such as mixed reviews and sarcasm. They usually improve consistency, especially for labels specific to your business.
Writing a good sentiment prompt
- Define each label precisely, including what "neutral" and "mixed" mean for you.
- Ask for structured output, such as JSON with a fixed set of allowed labels, so results are easy to parse.
- Ask for the key phrase behind the label. It helps debugging and audits, though a stated reason may not reflect how the model actually decided.
- Use a temperature of 0 or close to it for repeatable labels.
- Put the rules and examples first and the text to classify last.
You are a sentiment classifier for customer reviews of a food delivery app.
Labels:
- positive: the customer is satisfied overall
- negative: the customer is dissatisfied overall
- mixed: clear praise AND clear complaints about different things
- neutral: no opinion, only facts or questions
Return JSON only: {"label": "<one label>", "evidence": "<short quote from the review>"}
Examples:
Review: "Rider was polite but the food arrived cold."
-> {"label": "mixed", "evidence": "food arrived cold"}
Review: "Oh great, another 90 minute wait. Love it."
-> {"label": "negative", "evidence": "another 90 minute wait"}
Review: "{review_text}"Whatever provider you use, validate the output before trusting it:
import json
ALLOWED = {"positive", "negative", "mixed", "neutral"}
def classify(review: str, prompt: str, call_llm) -> dict:
# call_llm wraps your LLM client and sends the prompt with temperature 0
raw = call_llm(prompt.replace("{review_text}", review)) # replace(), not format(): the prompt contains JSON braces
try:
result = json.loads(raw)
except json.JSONDecodeError:
return {"label": "unknown", "evidence": raw[:200]}
if result.get("label") not in ALLOWED:
result["label"] = "unknown"
return resultWhere LLMs shine
- No labels yet, or label definitions that change often
- Nuanced judgments such as sarcasm, mixed opinions, aspect extraction and emotions with explanations
- Low volume, high value work, like analyzing a few hundred survey answers for a report
- Bootstrapping labels: let an LLM pre-label data, have people review a sample, then train a smaller, cheaper model on the result, an approach often called distillation
Where to be careful
- Cost and latency. Sending millions of short texts through a large model costs far more than running a fine-tuned small model.
- Consistency. Outputs can shift with prompt wording or model updates, so pin model versions and re-test after any change.
- Privacy. Sending customer text to a third-party API may need consent or contract terms, so mask personal data first.
- Evaluation. Measure the LLM against a hand-labeled test set, exactly as you would any other model.
Aspect-Based Sentiment Analysis
Overall sentiment tells you how customers feel; aspect-based sentiment analysis (ABSA) tells you why. It finds the specific things people talk about, called aspects, and the sentiment toward each one.
Take "The camera is stunning, but the battery barely lasts a day and support never replied." Document-level analysis might call this mixed or negative. ABSA returns three separate findings:
Aspect | Opinion words | Sentiment |
|---|---|---|
Camera | stunning | Positive |
Battery | barely lasts a day | Negative |
Customer support | never replied | Negative |
The sub-tasks
- Aspect term extraction finds the target words in the text, such as "battery" and "support."
- Aspect category detection maps them to a fixed list your team cares about, such as BATTERY, SERVICE and PRICE. It also catches implicit aspects: "way too expensive" is about PRICE without the word "price."
- Aspect sentiment classification decides the polarity toward each aspect.
- Opinion term extraction pulls out the words expressing each opinion, which makes results easy to explain.
The SemEval shared tasks of 2014 to 2016, built on restaurant and laptop reviews, defined many of the standard ABSA benchmarks.
Approaches
- Rules plus parsing. Find nouns as aspects and the adjectives or verbs attached to them, then score those words with a lexicon. Transparent, but brittle.
- Sentence-pair transformers. Feed the model the review and the aspect as a pair and classify the sentiment toward that aspect. This works well when you have a fixed category list.
- Tag, then classify. Mark aspect spans with an NER-style token classifier, then classify each span.
- End-to-end generation. Sequence-to-sequence models or LLMs output (aspect, category, opinion, sentiment) tuples directly, which is especially handy for prototyping.
A sentence-pair example with a community ABSA model:
from transformers import pipeline
absa = pipeline("text-classification", model="yangheng/deberta-v3-base-absa-v1.1")
review = "The camera is stunning, but the battery barely lasts a day."
for aspect in ["camera", "battery"]:
print(aspect, absa({"text": review, "text_pair": aspect}))Community models differ in input format and label names, so check the model card before relying on one.
Making ABSA useful
- Agree on categories with the people who will use them. Ten to twenty categories are usually more actionable than hundreds of raw aspect terms.
- Merge synonyms, so "battery," "charge" and "battery life" count as one aspect.
- Track aspects over time and by product version, so teams can see whether a fix actually changed how customers feel.
Evaluating Sentiment Models
A sentiment model is only useful if you know how well it works on your data. Metrics tell you how much it gets wrong, and error analysis tells you what to fix next.
Metrics that matter
- Accuracy is the share of correct predictions. It misleads on imbalanced data: if 85% of reviews are positive, a model that always says "positive" scores 85%.
- Precision asks how many texts predicted negative truly are negative. It matters when false alarms are costly, such as escalating calm customers.
- Recall asks how many truly negative texts the model caught. It matters when misses are costly, such as overlooking a customer about to leave.
- F1 is the harmonic mean of precision and recall.
- Macro F1 averages F1 across classes, treating each class equally. It is the best single number for imbalanced sentiment data, while weighted F1 can hide poor results on rare classes.
- Mean absolute error suits star ratings, since predicting 4 stars for a 5-star review is far better than predicting 1.
- Calibration checks whether a model that says "90% confident" is right about 90% of the time.
The confusion matrix
A confusion matrix shows which classes get mixed up. In three-class sentiment, most errors usually involve neutral, while outright positive-for-negative swaps are rarer and often point to sarcasm or label noise.
import matplotlib.pyplot as plt
from sklearn.metrics import ConfusionMatrixDisplay, classification_report
labels = ["negative", "neutral", "positive"]
print(classification_report(y_true, y_pred, labels=labels, digits=3))
ConfusionMatrixDisplay.from_predictions(y_true, y_pred, labels=labels)
plt.show()Error analysis
Pull 50 to 100 misclassified examples and tag each with a reason: sarcasm, negation, mixed sentiment, domain slang, missing context or a wrong label. Count the tags. The biggest bucket tells you what to do next, whether that is more data of one kind, better preprocessing, a clearer label definition or a stronger model.
Expect a surprising share of "errors" to be labeling mistakes. Fixing them improves both training and the honesty of your test scores.
Slice-based evaluation
Overall scores hide weak spots, so measure performance on slices: short versus long texts, each product line, each language, each channel and texts with emojis or negation. A model with strong overall accuracy can still do badly on Roman Urdu messages or on one product category.
Behavioral tests
The CheckList method (Ribeiro and colleagues, 2020) suggests small targeted tests:
- Invariance: changing a name or city should not change the label, so "Ali loved it" and "Sara loved it" must match.
- Directional: adding "not" should flip or weaken positive sentiment, and adding "absolutely" should strengthen it.
- Minimum functionality: simple sentences the model must always get right, such as "I hate it" being negative.
These tests catch regressions that aggregate metrics miss, and they make good automated checks before every new model release.
Compare against a baseline
Always report a simple baseline, such as VADER or TF-IDF with logistic regression, next to your model on the same test set. It shows whether the complex model earns its extra cost.
The Hard Problems: Sarcasm, Domain Shift, Multilingual Text and Bias
Even the best models still struggle with some kinds of text. Knowing these limits helps you set honest expectations and design around them.
Sarcasm and irony
"Wow, waited two hours for cold pizza. Amazing service." Every sentiment word is positive, yet the meaning is clearly negative. Sarcasm depends on the clash between words and situation, on shared context and on a tone of voice that text usually lacks.
Helpful signals include positive words describing bad events, exaggeration, scare quotes and eye-roll emojis. Transformers and LLMs catch obvious sarcasm far better than lexicons, but subtle cases still fool them, and human annotators disagree on these too. Track sarcasm as its own error category, and include sarcastic examples in both training and test data.
Implicit sentiment and comparisons
"The battery lasted two hours" contains no opinion words, yet for a phone it is plainly negative. "Better than the old version" is positive about one thing and negative about another, and "I expected more" is a complaint without a single negative word. These cases need world knowledge and comparison handling, which is where aspect-based methods and LLMs help most.
Domain shift
Models trained in one domain lose accuracy in another. "Long" is good for battery life and bad for wait times, and "cheap" depends on whether the topic is price or quality. Language also drifts over time, as new slang, product names and events change what words signal.
The fixes are practical. Fine-tune on in-domain data, even a few hundred examples, then monitor production performance and retrain on a schedule.
Multilingual and code-mixed text
Most sentiment resources are English-first, and lower-resource languages have fewer datasets and weaker lexicons. Code-mixed text such as "service bilkul bakwas thi" (Roman Urdu for "the service was total rubbish") blends vocabularies and has no standard spelling.
You have three main options:
- Multilingual models such as XLM-RoBERTa, fine-tuned on a small local dataset
- Translation to English before analysis, which is quick but loses slang and nuance
- LLMs, which often handle code-mixing well but must be tested on real samples first
Bias and fairness
Sentiment models can absorb social biases from their training data. The Equity Evaluation Corpus study (Kiritchenko and Mohammad, 2018) found many systems scored otherwise identical sentences differently depending on the gender or race associated with names in them. Studies have found similar problems with dialects, where some varieties of English get scored more negatively.
If your scores feed decisions about people, such as moderation or customer prioritization, test with counterfactual pairs. Swap names, genders and dialect markers, then compare results across groups.
Labels are partly subjective
Sentiment is partly in the eye of the reader. Annotators often disagree on mild or mixed texts, which caps the accuracy any model can reach. Rather than fight this, consider soft labels that record the share of annotators choosing each class, or a dedicated "uncertain" class, and report annotator agreement next to model scores.
Real-World Applications and Deploying to Production
A sentiment model creates value only when its output reaches a decision. This section covers common applications and what it takes to run a model reliably.
Where sentiment analysis is used
- E-commerce and app reviews: rank recurring complaints, flag reviews that need a reply, compare product versions.
- Customer support: prioritize angry or urgent tickets, track how sentiment changes from first message to resolution, coach agents.
- Social media monitoring: track brand mentions, catch sudden negative spikes, measure reactions to campaigns.
- Surveys and employee feedback: summarize thousands of open-ended answers by theme and sentiment.
- Finance: domain models such as FinBERT score sentiment in earnings calls, filings and news, used as one signal among many.
- Healthcare and public services: analyze patient or citizen feedback, under strict privacy controls.
Serving a model as an API
For web applications, wrap the model in a small API. Here is a FastAPI example that serves the model fine-tuned earlier:
from fastapi import FastAPI
from pydantic import BaseModel
from transformers import pipeline
app = FastAPI()
clf = pipeline("sentiment-analysis", model="./sentiment-model")
class Reviews(BaseModel):
texts: list[str]
@app.post("/sentiment")
def predict(req: Reviews):
results = clf(req.texts, truncation=True, batch_size=32)
return [{"text": t, "label": r["label"], "score": round(r["score"], 4)}
for t, r in zip(req.texts, results)]Run it with uvicorn app:app, and any backend, whether Laravel, Node.js or Django, can call the endpoint over HTTP.
Production checklist
- Batch and cache. Most sentiment workloads are throughput-bound, and repeated texts need not be scored twice.
- Size hardware to volume. Distilled models on CPUs handle moderate traffic; GPUs or quantized ONNX models handle more.
- Ship preprocessing with the model, so serving runs exactly what training ran.
- Log the model version with every prediction, so results can be traced and reprocessed.
- Monitor drift. Watch inputs (new vocabulary, languages, text lengths) and outputs (a sudden jump in the negative share). A spike may be a real crisis or a broken data feed, and dashboards should make it easy to tell which.
- Keep people in the loop for high-stakes actions, and route low-confidence predictions to review.
- Close the loop. Feed reviewer corrections into the next training round.
- Protect privacy. Mask personal data, set retention limits and follow local data protection law.
Presenting results
Aggregate with care. Always show volume beside the negative share, because a jump from 10% to 40% negative means little if it covers three messages. Link every chart to example texts so stakeholders can read what customers actually said, and combine sentiment with aspects to answer the question that matters most: what should we fix first?
Conclusion: Choosing the Right Approach
Sentiment analysis has never had more good options. The right one depends on your data, budget, speed requirements and how much you need to explain each decision.
Your situation | Good starting point |
|---|---|
No labeled data, results needed today | VADER, or an LLM with a clear prompt |
Some labeled data, need speed and explanations | TF-IDF with logistic regression or a linear SVM |
A few thousand labels, accuracy matters most | Fine-tuned DistilRoBERTa or RoBERTa |
Need to know what customers like and dislike | Aspect-based analysis with a fixed category list |
Many languages or code-mixed text | XLM-RoBERTa fine-tuned on local data, tested against an LLM |
Small volume, nuanced judgments | An LLM with few-shot examples and JSON output |
Whatever you choose, the same habits decide success. Define labels clearly, build a trustworthy test set from your own domain, compare against a simple baseline, read the errors, and keep monitoring after launch.
Next steps
Start small. Take 500 real texts from your own product or community, label them carefully, and compare VADER, a TF-IDF model, a ready-made transformer and an LLM prompt on that set. One afternoon will show which approach fits your data, and you will end up with the test set everything else depends on.
From there, explore aspect-based analysis to learn why people feel the way they do, and emotion detection to capture what positive and negative alone miss. Sentiment analysis is one of the oldest NLP tasks and still one of the most practical: done well, it turns what customers write into what teams do next.

Comments
Post a Comment