From the Hub and Transformers v5 to Fine-Tuning, Inference, Spaces and Agents
Updated October 2026 · 34 min read
Table of Contents
- Introduction: The Platform Behind Open AI Development
- The Modern AI Tech Stack in Seven Layers
- A Short History: From Chatbot App to AI Infrastructure
- The Hugging Face Hub: The Library of AI
- Getting Started: Account, Tokens and the hf CLI
- Transformers v5: The Model Definition Layer
- Tokenizers: The First and Last Step of Every Model
- Datasets: Data Loading Built for Machine Learning
- Sentence Transformers: Embeddings for Search and RAG
- Diffusers: Image, Video and Audio Generation
- The Fine-Tuning Stack: Trainer, Accelerate, PEFT and TRL
- Evaluation: Knowing Whether Your Model Is Good
- Inference and Deployment: Running Models for Real Users
- Spaces and Gradio: Shipping AI Demos and Apps
- Agents: smolagents, MCP and HuggingChat
- Beyond Text: Vision, Audio and Robotics
- Pricing and Plans in 2026
- Putting It All Together: A Reference Architecture
- Security, Licensing and Responsible Use
- A 30-Day Learning Roadmap
- Frequently Asked Questions
- Conclusion: The Open AI Stack Is Yours to Build With
Introduction: The Platform Behind Open AI Development
If you have downloaded an open AI model, fine-tuned a language model, or tried a viral AI demo in your browser, there is a very good chance you touched Hugging Face along the way, even if you did not notice. In less than a decade, a company that started as a chatbot app for teenagers has become the central meeting place of the open machine-learning world: the place where models are published, datasets are shared, demos are hosted and, increasingly, where production AI systems get their building blocks.
People often describe Hugging Face as “the GitHub of AI”, and the comparison is a good starting point. GitHub hosts code and lets developers collaborate on it; Hugging Face hosts models, datasets and AI apps and lets researchers, companies and hobbyists collaborate on them. But the analogy undersells it. GitHub does not also write the most popular programming libraries, run inference for you on dozens of cloud providers, or give you free GPUs for demos. Hugging Face does all of that.
A better analogy might be a city for AI builders. The Hub is the city’s library and marketplace, where millions of models and datasets sit on open shelves. The open-source libraries (Transformers, Datasets, Diffusers, PEFT, TRL and others) are the city’s workshops, full of standard tools anyone can use. Spaces are the city’s exhibition halls, where people show off what they built. Inference Providers and Endpoints are the city’s power plants and transport network, delivering compute where it is needed. And the community of researchers, startups and big tech companies is the population that keeps the whole place alive.
This guide is a complete tour of that city as it stands in 2026. We will start with the big picture of the modern AI tech stack, then walk through each Hugging Face component with clear explanations and working code: the Hub, the hf command-line tool, Transformers v5, Tokenizers, Datasets, Sentence Transformers, Diffusers, the fine-tuning stack, evaluation, the many ways to run inference, Spaces and Gradio, and the agent ecosystem. We finish with pricing, a reference architecture for a real application, a 30-day learning roadmap and answers to common questions.
Who This Guide Is For
- Developers who want to add AI features (chat, search, classification, image generation) to web or mobile apps using open models.
- Students and data scientists who want a clear map of the ecosystem before diving into tutorials.
- Technical founders and team leads deciding whether open models on Hugging Face can replace or complement closed AI APIs.
You need basic Python knowledge. Some examples benefit from a GPU, but nearly everything can be tried on a laptop, a free Google Colab notebook or a free Hugging Face Space.
Versions used in this guide
The examples target the 2026 generation of the ecosystem: Transformers v5, huggingface_hub v1.x with the hf command-line tool, and current releases of Datasets, PEFT, TRL, Diffusers and Gradio. Hugging Face now ships Transformers minor releases weekly, so check the documentation for the latest details when something behaves differently.
The Modern AI Tech Stack in Seven Layers
Before zooming into Hugging Face, it helps to see the whole AI tech stack from above. Every AI application, from a simple sentiment classifier to an autonomous agent, is built from the same layers. Each layer depends on the one below it.
| Layer | What it does | Typical examples | Hugging Face role |
|---|---|---|---|
| 7. Applications & agents | The product users interact with | Chatbots, copilots, search, agents | Spaces, Gradio, smolagents, HuggingChat |
| 6. Inference & serving | Running models efficiently for users | vLLM, SGLang, llama.cpp, cloud APIs | Inference Providers, Inference Endpoints, Transformers.js |
| 5. Training & adaptation | Teaching models new skills | Fine-tuning, LoRA, RLHF, distillation | Transformers, PEFT, TRL, Accelerate |
| 4. Models | The learned “brains” | LLMs, embedding, vision, speech, diffusion models | The Hub (millions of models), Transformers model definitions |
| 3. Data | What models learn from and are tested on | Web text, instructions, preferences, images | The Hub (datasets), Datasets library, Storage Buckets |
| 2. Frameworks | Math and automatic differentiation | PyTorch, JAX, MLX | Built on top of these (Transformers is now PyTorch-first) |
| 1. Compute | The hardware that does the work | GPUs, TPUs, CPUs, accelerators | ZeroGPU, Spaces hardware, Jobs, partner clouds |
The striking thing about this table is the right-hand column. Hugging Face has a presence in almost every layer above the raw hardware. That is why understanding Hugging Face is, in practice, a shortcut to understanding the open AI stack as a whole.
Open vs Closed: Where Hugging Face Fits
There are two broad ways to build with AI in 2026. The closed path uses proprietary models through an API: you send text, you get text back, and you never see the model weights. The open path uses models whose weights you can download, inspect, fine-tune and run wherever you like.
Neither path is universally better. Closed APIs are convenient and often lead on the very hardest tasks. Open models offer control, privacy, customisation, predictable costs at scale and freedom from vendor lock-in, and the best open models have become remarkably capable. Many real systems mix both: a closed model for complex reasoning, and open models for embeddings, classification, moderation or high-volume tasks.
Hugging Face is the home of the open path, but it also bridges the two. Its Inference Providers service offers an OpenAI-compatible API for open models hosted by many companies, so switching between closed and open models can be as simple as changing a URL and a model name.
A Short History: From Chatbot App to AI Infrastructure
Understanding where Hugging Face came from explains a lot about its culture and products.
- 2016 – A chatbot for teenagers. Hugging Face was founded in New York as a consumer app: a friendly AI companion with the hugging-face emoji as its logo. To build it, the team had to work with cutting-edge NLP research.
- 2018–2019 – The pivot to open source. When Google released BERT, the team published a PyTorch implementation that quickly became more popular than their app. That project grew into the Transformers library, which offered a single, consistent interface to many model architectures.
- 2020–2021 – The Hub is born. To share pretrained weights, Hugging Face built the Model Hub, then added Datasets and Spaces. Sharing a model became as easy as pushing a Git repository.
- 2022 – Open science at scale. Hugging Face coordinated BigScience, a year-long collaboration of over a thousand researchers that produced the BLOOM language model, and the open-source community exploded around Stable Diffusion and its Diffusers library.
- 2023–2024 – The open LLM boom. Llama, Mistral, Qwen, Gemma, DeepSeek and many other model families were released on the Hub. Hugging Face added tools for efficient fine-tuning (PEFT, TRL), inference, and evaluation, and its user base grew into the millions.
- 2025 – Platform maturity. The company released huggingface_hub v1.0 with the new
hfcommand-line tool, launched Inference Providers to route requests across many serverless providers, moved storage onto the Xet chunk-based backend, expanded into open-source robotics with LeRobot, and announced Transformers v5, the first major version in five years. - 2026 – Infrastructure for the open ecosystem. Transformers v5 shipped with weekly releases, Hugging Face introduced S3-like Storage Buckets on the Hub in March 2026, and its libraries became the shared model-definition layer used by inference engines like vLLM, SGLang and llama.cpp.
The lesson from this history is that Hugging Face succeeds by being the neutral place where everyone shares. Big tech companies, startups, universities and individuals all publish there because that is where users are, and users come because that is where the models are.
The Hugging Face Hub: The Library of AI
The Hub at huggingface.co is the heart of the ecosystem. Technically, it is a collection of Git-based repositories with a rich website on top. Each repository is one of three main types: a model, a dataset or a Space. Since 2026, there is also a fourth storage type, Buckets, for files that do not need version control.
Models
The Hub hosts well over two million public model repositories, covering text generation, embeddings, translation, speech recognition, text-to-speech, image classification, object detection, image and video generation, robotics policies and much more. Each model repository typically contains:
- Weights, usually in the
safetensorsformat (a safe, fast format that cannot execute hidden code, unlike older pickle files). - A configuration file (
config.json) describing the architecture. - Tokenizer or processor files that define how raw input becomes numbers.
- A model card (
README.md) explaining what the model does, how it was trained, its limitations, intended uses and license.
Model popularity is extremely concentrated. A small number of base models and their derivatives account for most downloads, while a long tail of fine-tunes serves specific languages, domains and tasks. That long tail is a hidden treasure: if you need a model for Urdu sentiment, legal document classification or medical entity recognition, someone may already have fine-tuned one.
How to Read a Model Page Like a Pro
When you open a model page, check these elements before using it:
| Element | What to look for |
|---|---|
| License | Apache-2.0 and MIT are permissive. Some families use custom licenses with usage restrictions; read them before commercial use. |
| Gated access | Some models require you to accept terms and request access before downloading. |
| Model card quality | Good cards describe training data, evaluation results, limitations and intended use. A blank card is a warning sign. |
| Downloads and likes | Popularity is not quality, but it signals community testing. |
| Model tree | Shows whether this is a base model, a fine-tune, an adapter, a quantization or a merge, and links to its parent. |
| Files and versions | Check file sizes (memory needed), formats (safetensors, GGUF) and commit history. |
| Inference widget and “Use this model” | Try the model in the browser and copy ready-made code for libraries and local apps. |
Pin a revision for production
Model repositories can change. For reproducible applications, pass a specific commit hash with revision="..." when loading a model, so an upstream update cannot silently change your results.
Datasets
The Hub also hosts hundreds of thousands of datasets: instruction-tuning data, preference pairs, benchmark suites, speech corpora, image collections and domain-specific text. Datasets are stored in repositories just like models, very often as Parquet files, and each dataset page includes a viewer that lets you browse rows, search and filter in the browser, along with a dataset card describing its origin and license.
Spaces
Spaces are hosted AI applications. A Space is a repository containing an app written with Gradio, Streamlit, static HTML or any framework packaged in Docker. Push code, and Hugging Face builds and runs it at a public URL. Spaces range from small demos to sophisticated tools, and they are the fastest way to share a working AI app with the world, or with your boss.
Storage Buckets
Introduced in March 2026, Storage Buckets provide mutable, S3-like object storage on the Hub. Model and dataset repositories are versioned with Git, which is perfect for final, stable artifacts but awkward for files that change constantly, such as training checkpoints, optimizer states, logs and intermediate processed data. Buckets are not versioned, are built on the Xet storage backend (which deduplicates content across similar files), and are managed with the hf buckets commands. A good mental model: repos for what you publish, buckets for what you are still working on.
Organisations, Collections and Community Features
- Organisations let teams and companies share repositories, manage permissions and pay for compute together.
- Collections are curated lists of models, datasets, Spaces and papers, useful for following a model family or building a reading list.
- Discussions and pull requests on every repository let users report issues, ask questions and propose fixes.
- Daily Papers is a community-curated feed of new AI research papers, each linked to related models, datasets and demos.
The Xet Storage Backend
Model files are huge, and many fine-tunes differ from their parents by only small amounts. Hugging Face moved repository storage from Git LFS to Xet, a chunk-based system that splits files into pieces and stores identical pieces only once. For users, the main effects are faster uploads and downloads of large files, especially when only part of a file changes. You mostly do not need to think about it: the huggingface_hub library handles it automatically.
Getting Started: Account, Tokens and the hf CLI
You can download most public models without an account, but creating a free account unlocks uploading, gated models, Spaces, inference credits and higher rate limits.
Step 1: Create an Access Token
After signing up, open Settings → Access Tokens and create a token. Hugging Face offers fine-grained tokens, which let you limit a token to specific repositories and permissions (for example, read-only access to gated models, or write access to one organisation). Follow the principle of least privilege: a token for a deployed app should only be able to do what that app needs.
Never commit tokens
Store tokens in environment variables (HF_TOKEN) or a secrets manager, never in code or notebooks you share. Hugging Face scans public repositories for leaked tokens, but prevention is far better than revocation.
Step 2: Install the Libraries and Log In
The huggingface_hub library is the Python client for everything on the Hub. Version 1.0, released in 2025, introduced the modern hf command-line tool, which replaced the older huggingface-cli command.
Step 3: Download and Upload with the CLI
The CLI covers the everyday operations you would otherwise click through on the website:
Downloads go into a shared cache (by default under ~/.cache/huggingface). Every library in the ecosystem uses the same cache, so a model downloaded once by the CLI is instantly available to Transformers, Diffusers and others.
Step 4: The Same Operations in Python
The HfApi class exposes almost everything on the Hub: searching, creating and deleting repositories, managing branches and pull requests, reading model metadata and more. It is the foundation for automation such as nightly model uploads from a training pipeline.
Speed up big downloads
For multi-gigabyte models, the default parallel downloader is usually fast, but check the huggingface_hub documentation for current environment variables that tune concurrency on very fast connections. Also set HF_HOME to a disk with plenty of free space before downloading large models.
Transformers v5: The Model Definition Layer
Transformers is Hugging Face’s flagship library and one of the most widely used Python packages in the world. Its job is to give you a consistent way to load, run and train hundreds of model architectures for text, vision, audio and multimodal tasks.
What Changed in v5
Transformers v5, announced at the end of 2025, was the first major version in five years. By then, the library was being installed more than three million times per day, supported over 400 model architectures, and was compatible with more than 750,000 model checkpoints on the Hub. Version 5 was less about flashy new features and more about becoming solid shared infrastructure. The highlights:
- Simpler, more modular model code. Model definitions were refactored to reduce duplication, making it easier to add new architectures and to read existing ones.
- PyTorch-first. v5 went “all in” on PyTorch as its framework, dropping the parallel TensorFlow and JAX implementations of earlier versions.
- A shared reference for the ecosystem. Inference engines such as vLLM, SGLang and llama.cpp, and training frameworks, increasingly rely on Transformers model definitions rather than reimplementing every architecture themselves.
- Better inference built in. Specialised kernels, cleaner defaults, continuous batching and paged attention arrived for serving use cases.
- Faster release cadence. After v5.0, minor releases moved to a weekly schedule, so new model architectures land quickly.
- Cleanup of legacy APIs. Older patterns were removed or deprecated, such as some Text2Text pipelines and legacy chat-template saving. A migration guide covers the breaking changes.
Pipelines: AI in Three Lines
The pipeline() function is the easiest door into Transformers. It bundles a model, its preprocessing and its postprocessing into a single callable object:
Pipelines exist for dozens of tasks. A few of the most useful:
| Task name | What it does |
|---|---|
text-generation | Chat and text completion with LLMs |
text-classification / sentiment-analysis | Label a text with a category |
token-classification / ner | Find entities such as names and places |
zero-shot-classification | Classify into labels you choose at runtime, without training |
feature-extraction | Produce embeddings from text |
automatic-speech-recognition | Transcribe audio (for example with Whisper models) |
image-classification, object-detection | Understand images |
image-text-to-text | Ask questions about images with vision-language models |
Chatting with an Open LLM
Modern pipelines accept chat messages directly and apply the model’s chat template for you:
device_map="auto" places the model on your GPU if you have one (and can split very large models across several devices), falling back to the CPU otherwise.
We use SmolLM2, a small open model trained by Hugging Face, because it runs on almost any laptop. Some newer model families, such as Qwen3, are “hybrid reasoning” models that write out a thinking section before answering by default; their model cards explain how to switch that behaviour on or off through the chat template.
Under the Hood: Auto Classes, Tokenizers and Chat Templates
Pipelines are convenient, but real applications often need more control. The Auto classes load the right architecture from a model’s configuration automatically:
Three ideas in this snippet are worth understanding deeply:
- The tokenizer converts text into token IDs and back. Every model has its own tokenizer, and using the wrong one produces nonsense.
- The chat template is a small program stored with the tokenizer that formats a list of messages into the exact string format the model was trained on, including special tokens marking where the user’s turn ends and the assistant’s begins. Getting this format wrong is one of the most common reasons an open model “seems dumb”.
- Generation parameters control the output.
max_new_tokenslimits length;temperatureanddo_samplecontrol randomness. Low temperature gives focused, repeatable answers; higher temperature gives more varied, creative ones.
Running Bigger Models on Smaller Hardware: Quantization
A model’s memory footprint is roughly its number of parameters multiplied by the bytes per parameter. A 7-billion-parameter model in 16-bit precision needs about 14 GB just for its weights. Quantization stores weights with fewer bits (8-bit, 4-bit or even less), cutting memory dramatically with a modest quality cost.
| Precision | Bytes per parameter | Memory for a 7B model (weights only) |
|---|---|---|
| 32-bit float | 4 | ~28 GB |
| 16-bit (bf16/fp16) | 2 | ~14 GB |
| 8-bit | 1 | ~7 GB |
| 4-bit | 0.5 | ~3.5 GB |
Transformers supports several quantization backends. The classic beginner-friendly option is bitsandbytes 4-bit loading:
The Hub also hosts many pre-quantized models in formats such as GGUF (for llama.cpp and Ollama), AWQ, GPTQ and others. Searching for a model name plus “GGUF” usually finds ready-to-run versions.
Serving a Model from Transformers
For quick experiments, Transformers can expose a model through a local OpenAI-compatible server with the transformers serve command, so tools that speak the OpenAI API can talk to your local model. For production traffic, dedicated engines like vLLM and SGLang (covered in the inference section) are faster, and they build on the same Transformers model definitions.
Tokenizers: The First and Last Step of Every Model
The Tokenizers library is a fast, Rust-based implementation of the tokenization algorithms used by modern models: Byte-Pair Encoding (BPE), WordPiece and Unigram. You rarely call it directly, because AutoTokenizer uses it automatically, but it matters for three reasons.
First, speed. Tokenizing millions of documents for training or analysis is fast because the heavy lifting happens in compiled code with automatic batching.
Second, cost and context. LLM context windows and API bills are measured in tokens, not words. Counting tokens with the right tokenizer tells you exactly how much text fits in a prompt:
Third, languages. Tokenizers trained mostly on English split other languages into many more tokens. The same sentence in Urdu, Hindi or Arabic can cost several times more tokens than in English with some tokenizers, which means higher costs, smaller effective context windows and sometimes weaker performance. When choosing a model for a non-English product, compare token counts on your own text: it is a quick, revealing test.
You can also train your own tokenizer on a domain corpus (for example, code or legal text) in minutes with this library, which is a standard step when pretraining a model from scratch.
Datasets: Data Loading Built for Machine Learning
The Datasets library loads and processes data for machine learning efficiently. Its secret is Apache Arrow: datasets are stored on disk in a columnar format and memory-mapped, so you can work with datasets far larger than your RAM as if they were in memory.
Loading Data from the Hub or Your Files
Transforming Data with map()
The map() method applies a function to every example, and with batched=True it processes many rows at once, which is dramatically faster, especially for tokenization:
Results are cached automatically. If you rerun the same map() with the same function and inputs, Datasets reuses the cached result instead of recomputing it.
Streaming Huge Datasets
Some datasets on the Hub are measured in terabytes. With streaming=True, you iterate over examples as they download, without storing the whole dataset:
Sharing Your Own Dataset
Publishing a dataset with a good dataset card (source, collection method, license, known biases and intended use) is one of the most valuable contributions you can make to the community, and a private dataset repository is an excellent way to version data for your own team.
Sentence Transformers: Embeddings for Search and RAG
Sentence Transformers is the go-to library for embeddings: dense vectors that represent the meaning of text (and, with multimodal models, images). Embeddings power semantic search, clustering, recommendation, duplicate detection and, most importantly in 2026, Retrieval-Augmented Generation (RAG), where an LLM answers questions using documents retrieved by meaning.
The query shares no keywords with the refund sentence, yet it ranks first, because embeddings capture meaning rather than exact words. That is the core magic of semantic search.
Choosing an Embedding Model
The Hub hosts thousands of embedding models. When choosing one, consider:
- Quality on your task and languages. The community-maintained MTEB leaderboard compares embedding models across many tasks and languages; use it as a starting shortlist, then test on your own data.
- Size and speed. Small models embed thousands of sentences per second on a CPU; large models are more accurate but slower and more expensive.
- Dimension. Smaller vectors mean cheaper storage and faster search. Some models support “Matryoshka” embeddings that can be truncated to fewer dimensions with little quality loss.
- Multilingual support. For products serving users in several languages, pick a model trained on multilingual data.
Sentence Transformers also supports rerankers (cross-encoders), which read a query and a document together for a more precise relevance score. A common high-quality pattern is to retrieve the top 50 candidates with fast embeddings, then rerank them with a cross-encoder and pass the best few to the LLM.
Diffusers: Image, Video and Audio Generation
Diffusers is Hugging Face’s library for diffusion models, the family behind most modern image and video generators. Diffusion models learn to create data by reversing a gradual noising process: starting from pure noise, they remove a little noise at each step, guided by your text prompt, until an image emerges.
Diffusers offers modular building blocks (pipelines, models, schedulers) and techniques for running large generators on modest hardware, such as CPU offloading and quantization. It also supports LoRA adapters for image models, which the community uses to teach a model a specific style, character or product appearance with only a few dozen training images.
Mind the license
Image and video model licenses vary widely. Some allow commercial use, others restrict it, and some distilled “fast” variants have different terms from their larger siblings. Always check the model card before using generated images in a product.
The Fine-Tuning Stack: Trainer, Accelerate, PEFT and TRL
Pretrained models are generalists. Fine-tuning turns a generalist into a specialist: a support bot that knows your product, a classifier for your categories, or a model that writes in your company’s tone. Hugging Face provides a layered set of libraries for this, and understanding how they fit together removes most of the confusion beginners feel.
| Library | Role | Analogy |
|---|---|---|
| Trainer (in Transformers) | A ready-made training loop with logging, checkpointing and evaluation | The car’s automatic gearbox |
| Accelerate | Runs the same training code on one GPU, many GPUs or many machines | The road network connecting cities |
| PEFT | Trains small adapter weights instead of the whole model (LoRA and friends) | Adding a specialised attachment instead of rebuilding the machine |
| TRL | Trainers for LLM post-training: SFT, preference optimisation, reinforcement learning | A driving school for language models |
| bitsandbytes | Low-bit quantization so big models fit in small GPUs | Vacuum-packing luggage |
When Should You Fine-Tune at All?
Fine-tuning is powerful, but it is not always the first tool to reach for. A practical decision order:
- Prompting first. A clear system prompt with a few examples solves many tasks with zero training.
- RAG second. If the model lacks knowledge (your documents, policies, product data), retrieval usually beats fine-tuning, because you can update documents without retraining.
- Fine-tuning third. Fine-tune when you need a consistent behaviour, format or style, a smaller and cheaper model that matches a larger one on a narrow task, or a skill that prompting cannot reliably produce.
Classic Fine-Tuning with Trainer
For classification and other “encoder” tasks, the Trainer API is concise and battle-tested:
With push_to_hub=True, the trained model, tokenizer and an auto-generated model card are uploaded to your account, ready to share or deploy.
LoRA: Fine-Tuning Without Retraining Everything
Full fine-tuning updates every parameter of a model, which for an 8-billion-parameter LLM requires many high-end GPUs. LoRA (Low-Rank Adaptation) freezes the original weights and trains small “adapter” matrices added alongside certain layers. The adapters are often less than one percent of the model’s size.
Imagine a huge, expensive encyclopedia. Instead of reprinting it to add your company’s knowledge, you attach a slim booklet of sticky notes to the relevant pages. The encyclopedia stays the same; the notes change how it is read. That booklet is a LoRA adapter, and you can keep many different booklets for different tasks and swap them in seconds.
Combine LoRA with 4-bit quantization and you get QLoRA, which made it possible to fine-tune multi-billion-parameter models on a single consumer GPU or a free cloud notebook.
Post-Training LLMs with TRL
TRL (Transformer Reinforcement Learning) is the library for the post-training stages that turn a raw language model into a helpful assistant. Its trainers follow the same lifecycle used by AI labs:
- Supervised Fine-Tuning (SFT) teaches the model to follow instructions using example conversations.
- Preference optimisation (for example DPO) teaches it which of two answers humans prefer, without a separate reward model.
- Reinforcement learning (for example GRPO) improves it on tasks with checkable answers, such as maths or code, by rewarding correct outputs. This family of methods drove many of the “reasoning model” improvements of 2025.
A complete SFT run with LoRA is surprisingly short:
TRL reads conversational datasets in the standard messages format and applies the model’s chat template automatically. Swapping SFTTrainer for DPOTrainer or GRPOTrainer (with suitable datasets or reward functions) moves you to the next stage of post-training.
Data quality beats data quantity
A few thousand clean, diverse, carefully checked examples usually beat a hundred thousand noisy ones. Spend your time curating data: remove duplicates, fix wrong answers, balance topics and include the edge cases your users will actually hit.
Accelerate: From One GPU to Many
Accelerate lets the same training script run on a laptop, a single GPU, several GPUs or a cluster with minimal changes. Run accelerate config once to describe your hardware, then launch with accelerate launch train.py. Under the hood, it handles device placement, mixed precision and distributed strategies such as data parallelism and sharded training. Trainer and TRL use Accelerate automatically.
No GPU? Hugging Face Jobs
PRO users and paying organisations can run compute jobs on Hugging Face hardware with the hf jobs command, for example a fine-tuning script or a batch data-processing task, without managing servers. Combined with Storage Buckets for checkpoints and the Hub for final models, this creates an end-to-end training workflow entirely inside the platform.
Evaluation: Knowing Whether Your Model Is Good
Training without evaluation is guesswork. The open ecosystem offers several layers of evaluation, and serious teams use more than one.
Public Benchmarks and Leaderboards
Public leaderboards rank models on standardised benchmarks. They are useful for building a shortlist, but treat them with caution: popular benchmarks can leak into training data, models can be tuned to the test, and a high score on general knowledge says little about your specific task. Hugging Face retired its original Open LLM Leaderboard in 2025, partly because benchmark saturation made scores less meaningful, and the community now relies on a mix of specialised leaderboards (such as MTEB for embeddings) and task-specific evaluations.
Evaluation Libraries
- lighteval is Hugging Face’s toolkit for evaluating LLMs on many benchmarks with various backends, useful for reproducible comparisons.
- evaluate provides standard metrics (accuracy, F1, BLEU, ROUGE and others) with a simple interface.
Your Own Evaluation Set: The Most Important Benchmark
The best evaluation is a private test set built from your own use case: 100 to 500 real examples with known good answers, covering normal cases and tricky edge cases. Keep it out of training data, run every candidate model against it, and track results over time. For open-ended outputs, combine automatic checks (format, length, keywords, exact answers) with “LLM-as-a-judge” scoring and regular human review.
Inference and Deployment: Running Models for Real Users
Training gets the headlines, but inference (running a model to produce outputs) is where most money is spent and where user experience is decided. Hugging Face offers, and integrates with, many inference options. Choosing between them is one of the most practical skills in the 2026 AI stack.
Option 1: Local Inference on Your Own Machine
Running models locally gives you privacy, zero per-request cost and offline capability.
- Transformers runs any supported model in Python, ideal for experimentation.
- llama.cpp runs quantized GGUF models efficiently on CPUs, laptops and consumer GPUs. It can pull models directly from the Hub, for example
llama-server -hf ggml-org/gemma-3-1b-it-GGUF. - Ollama and LM Studio wrap local inference in a friendly app. Ollama can run GGUF models straight from the Hub with names like
hf.co/username/repo. - MLX runs models efficiently on Apple Silicon Macs.
Option 2: Self-Hosted Inference Servers
For production traffic on your own GPUs, use a dedicated serving engine:
- vLLM is a high-throughput LLM server known for PagedAttention memory management and continuous batching.
vllm serve Qwen/Qwen3-8Bstarts an OpenAI-compatible API. - SGLang is another high-performance engine, strong at structured outputs and complex multi-call workloads.
Hugging Face’s own Text Generation Inference (TGI) pioneered many of these serving ideas, but it has entered maintenance mode: it now accepts only minor fixes, and Hugging Face recommends vLLM, SGLang and local engines such as llama.cpp and MLX going forward. If you are starting a new project in 2026, choose one of those. If you already run TGI, plan a migration; the OpenAI-compatible APIs make it straightforward.
Option 3: Inference Providers (Serverless)
Inference Providers give you one API and one Hugging Face token to access thousands of models hosted by many serverless partners, including Cerebras, Groq, Together, Fireworks, Novita, Nscale, Replicate, fal, Featherless, Scaleway, OVHcloud and others, plus Hugging Face’s own HF Inference service. There is a free monthly allowance of credits, with more for PRO and enterprise accounts, and providers’ prices are passed through without extra markup.
The API is OpenAI-compatible, so existing code needs only a new base URL:
The suffix after the colon controls routing: :fastest picks the highest-throughput provider, :cheapest the lowest cost per token, and :groq or another provider name pins a specific provider. The huggingface_hub library offers the same capability through its InferenceClient, which also covers non-chat tasks such as text-to-image, speech recognition and embeddings.
Option 4: Inference Endpoints (Dedicated)
Inference Endpoints deploy any model from the Hub onto dedicated, autoscaling infrastructure in a few clicks, using engines such as vLLM, SGLang, llama.cpp or Text Embeddings Inference under the hood. Pricing is per hour of compute, from small CPU instances to multi-GPU machines, and endpoints can scale to zero when idle. Choose them when you need a private deployment of a specific model (including your own fine-tunes), predictable latency or compliance controls.
Option 5: In the Browser and at the Edge
- Transformers.js runs models directly in the browser or Node.js using WebGPU or WebAssembly. No server, no API cost, and user data never leaves the device:
- Optimum helps export and optimise models for specific hardware and runtimes such as ONNX Runtime, useful for mobile and edge deployment.
Choosing the Right Inference Option
| Your situation | Best starting option |
|---|---|
| Prototyping, learning, notebooks | Transformers locally or Inference Providers |
| Private data, offline use, laptops | llama.cpp, Ollama, LM Studio, MLX |
| Production app, variable traffic, no infrastructure team | Inference Providers |
| Production with your own fine-tuned model | Inference Endpoints, or vLLM/SGLang on your cloud |
| Very high volume on your own GPUs | vLLM or SGLang, self-hosted |
| Privacy-first web features | Transformers.js in the browser |
Spaces and Gradio: Shipping AI Demos and Apps
A model nobody can try is a model nobody uses. Spaces solve that by hosting AI apps for free on basic CPU hardware, with paid upgrades to GPUs when you need them.
Gradio in Five Minutes
Gradio, maintained by Hugging Face, builds a web interface around any Python function:
Save this as app.py, add a requirements.txt listing transformers and torch, push both files to a new Gradio Space, and within minutes the world can use your summariser at a public URL.
ZeroGPU: Free GPU Power for Demos
ZeroGPU is shared GPU infrastructure for Spaces. Instead of reserving a GPU all day, a ZeroGPU Space borrows a powerful GPU only while a function decorated with @spaces.GPU runs, then releases it. Free users can use ZeroGPU Spaces built by others within daily quotas, and PRO users get a much larger quota, higher queue priority and the ability to host their own ZeroGPU Spaces. It is one of the most generous offers in the AI world for students and indie developers.
Docker Spaces for Anything Else
If your app uses Node.js, FastAPI, a database or any custom stack, choose a Docker Space: provide a Dockerfile, and Hugging Face builds and runs the container. This makes Spaces suitable for full-stack AI prototypes, not just Python demos.
Gradio Apps as MCP Servers
The Model Context Protocol (MCP) is an open standard that lets AI assistants call external tools. Gradio can expose any app as an MCP server with a single argument:
Gradio turns each function into a tool, using its type hints and docstring as the tool description. That means every Space can become a capability that AI assistants and agents can use, which is why the docstring in the example above matters.
Agents: smolagents, MCP and HuggingChat
The biggest shift in applied AI since 2024 has been the move from models that answer to agents that act: they plan, call tools, read results and decide what to do next. Hugging Face contributes to this layer with a lightweight agent library, deep support for the Model Context Protocol and a free chat interface for open models.
smolagents: Agents That Think in Code
smolagents is Hugging Face’s minimalist library for building agents. Its distinctive idea is the CodeAgent: instead of asking the model to output tool calls as JSON, the agent writes small snippets of Python that call tools, combine results, loop and branch. Code is a more expressive language for actions than JSON, so code agents often solve multi-step tasks in fewer steps.
Notice how the @tool decorator turns an ordinary Python function into a tool. The type hints and the docstring, including the Args: section, become the tool description the model reads, so write them as if you were explaining the function to a new colleague.
Run code agents in a sandbox
A CodeAgent executes model-written code. smolagents restricts imports by default, but for anything beyond experiments, run agents inside a sandboxed environment (smolagents supports remote executors such as Docker-based and cloud sandboxes) and never give an agent credentials it does not strictly need.
MCP: The USB-C Port for AI Tools
The Model Context Protocol (MCP) standardises how AI applications discover and call tools and data sources. Hugging Face supports it at several levels:
- Gradio apps become MCP servers with
launch(mcp_server=True), so any Space can be a tool. - The Hugging Face MCP server lets compatible AI assistants search models, datasets and papers on the Hub and use community Spaces as tools.
- smolagents can use MCP tools, so an agent you build can plug into the same tool ecosystem as commercial assistants.
The result is a powerful network effect: every useful Space published by the community is potentially a capability for every MCP-compatible agent.
HuggingChat: Free Chat With Open Models
HuggingChat is Hugging Face’s chat interface for open models. Its “Omni” router, introduced in late 2025, automatically picks a suitable open model for each request from a pool of more than a hundred, so you get a good default without knowing which model is best at what. You can also select a specific model manually. For developers it doubles as a quick way to compare open models on your own prompts before committing to one.
Beyond Text: Vision, Audio and Robotics
Although Hugging Face started with NLP, the Hub now spans nearly every modality.
- Vision: image classification, object detection, segmentation and, increasingly, vision-language models that answer questions about images and documents. These are especially useful for reading invoices, forms and screenshots.
- Audio: speech recognition (Whisper-family models and many newer alternatives), text-to-speech, voice cloning and audio classification.
- Video: text-to-video and image-to-video generation through Diffusers, and video understanding through multimodal models.
- Robotics: LeRobot is Hugging Face’s open-source robotics library, providing datasets, pretrained policies and tools for training robots from demonstrations. After acquiring the robotics company Pollen Robotics in 2025, Hugging Face also began offering affordable open-source robot hardware, bringing the “open AI” approach to the physical world.
Pricing and Plans in 2026
One of Hugging Face’s strengths is how much you can do for free. Here is an overview of the main plans at the time of writing. Prices change, so confirm on the official pricing page before budgeting.
| Plan | Price | Best for | Key extras |
|---|---|---|---|
| Free | $0 | Learning, public projects | Unlimited public models, datasets and Spaces; free CPU Spaces; limited ZeroGPU and inference credits |
| PRO | $9 per month | Individual builders | Much more private storage, larger monthly inference credits, 8× ZeroGPU quota with top queue priority, host your own ZeroGPU Spaces, Spaces dev mode, Jobs |
| Team | $20 per user per month | Startups and teams | SSO, audit logs, resource groups for access control, PRO benefits for every member |
| Enterprise | From $50 per user per month | Large organisations | Highest limits, advanced security, user management, dedicated support and compliance help |
Compute is billed separately when you need it. Spaces run free on basic CPUs and can be upgraded to GPUs billed by the hour. Inference Endpoints start at a few cents per hour for small CPU instances, with GPU instances ranging from well under a dollar per hour for entry-level cards to many dollars per hour for multi-GPU servers. Extra storage is priced per terabyte per month, with lower prices for public data and volume discounts.
Students and indie developers
Before paying for any GPU, exhaust the free options: free CPU Spaces, community ZeroGPU Spaces, free monthly inference credits, and free notebook services. Upgrade to PRO when you hit the limits regularly; at $9 per month it is one of the cheapest ways to get real GPU access for demos.
Putting It All Together: A Reference Architecture
Let us design a realistic 2026 application using the Hugging Face stack end to end: a customer-support assistant that answers questions from a company’s help-centre articles, in more than one language, inside a web app.
The Architecture at a Glance
| Stage | Component | Hugging Face piece |
|---|---|---|
| 1. Store knowledge | Help articles as a versioned dataset | Private dataset repo, Datasets library |
| 2. Prepare | Split articles into chunks of a few hundred words | Datasets map() |
| 3. Embed | Turn chunks into vectors | Sentence Transformers (multilingual model) |
| 4. Index | Store vectors for fast search | FAISS, pgvector or a vector database |
| 5. Retrieve & rerank | Find the best chunks for each question | Embeddings + cross-encoder reranker |
| 6. Generate | Write an answer grounded in retrieved chunks | Open LLM via Inference Providers |
| 7. Serve | Web UI and API | Gradio on Spaces, or your own frontend |
| 8. Evaluate & improve | Measure accuracy, then fine-tune if needed | Private eval set, TRL + PEFT, Inference Endpoints |
The Core Retrieval and Generation Code
This short function captures the essence of modern RAG: fast semantic retrieval, precise reranking and a grounded generation prompt that tells the model to admit when it does not know. Note that the English-trained reranker here is a placeholder; for a truly multilingual product, choose a multilingual reranker from the Hub and verify it on your languages.
Calling It From a Web App
Because Inference Providers is OpenAI-compatible, a JavaScript or TypeScript frontend (for example a Next.js API route) can call open models with plain fetch:
Keep this call on the server side so your token is never exposed to browsers. The same HTTP request works from PHP (Laravel’s HTTP client), Go, Java or any language that can send JSON.
How the System Evolves
- Week 1: Ship the RAG version above behind a Gradio demo for internal testing.
- Month 1: Build a private evaluation set from real user questions; tune chunking, retrieval depth and prompts based on failures.
- Month 3: If costs or latency grow, fine-tune a small open model with TRL and LoRA on high-quality question–answer pairs, evaluate it against the large model, and deploy it on an Inference Endpoint or your own vLLM server.
- Ongoing: Version every dataset and model on the Hub, so you can always reproduce and roll back.
Security, Licensing and Responsible Use
Open models give you freedom, and freedom comes with responsibility. A professional checklist:
- Prefer safetensors. Older model formats based on Python pickle can execute arbitrary code when loaded. Safetensors cannot. Be cautious with repositories that only offer pickle files and with any loading option that asks you to trust remote code; read that code first.
- Read licenses carefully. Permissive (Apache-2.0, MIT), community licenses with conditions, and non-commercial licenses all exist on the Hub. The license of a fine-tune usually inherits obligations from its base model.
- Respect dataset terms and privacy. Know where your training data came from, remove personal information, and honour opt-outs.
- Pin versions and revisions of models and libraries in production.
- Protect tokens with fine-grained permissions, environment variables and regular rotation.
- Evaluate for safety, not just accuracy. Test how your application handles harmful requests, prompt injection in retrieved documents, and questions it should refuse.
- Document your work. Write model and dataset cards for anything you publish: intended use, limitations, evaluation results and known biases.
A 30-Day Learning Roadmap
If you are new to the ecosystem, this roadmap turns the guide into a plan.
| Week | Focus | Practical goal |
|---|---|---|
| 1 | Hub, hf CLI, pipelines | Run five different pipelines (sentiment, NER, chat, speech, image) and publish a Gradio demo on a free Space |
| 2 | Tokenizers, Datasets, embeddings | Load a Hub dataset, clean it with map(), build a semantic search over it with Sentence Transformers |
| 3 | Fine-tuning | Fine-tune a classifier with Trainer, then an LLM with TRL + LoRA on a small instruction dataset; push both to the Hub |
| 4 | Inference and agents | Call models via Inference Providers, run a GGUF model locally, build a smolagents agent with one custom tool, and expose a Space as an MCP server |
Hugging Face also publishes free courses on its website covering LLMs, agents, diffusion models, audio and more. They pair perfectly with this roadmap.
Frequently Asked Questions
Is Hugging Face free to use?
Yes, for most purposes. Browsing and downloading public models and datasets, publishing public repositories, hosting CPU Spaces and using limited ZeroGPU and inference credits are all free. You pay for private storage beyond the free limits, extra compute, dedicated endpoints and team or enterprise features.
Is Hugging Face only for NLP?
No. It began with NLP, but today the Hub covers computer vision, audio, video, multimodal models, time series, reinforcement learning and robotics.
What is the difference between Inference Providers and Inference Endpoints?
Inference Providers is serverless: you call popular models hosted by partner companies and pay per use, with no setup. Inference Endpoints are dedicated: you deploy a specific model, including private fine-tunes, on your own reserved hardware and pay per hour.
Do I need a GPU to use Hugging Face?
No. Many models run on CPUs, small models run on laptops, quantized models run through llama.cpp or Ollama, and you can use hosted inference, free Spaces and free notebooks for heavier work. A GPU helps for training and large models.
Is Text Generation Inference (TGI) still recommended?
TGI is now in maintenance mode, receiving only minor fixes. For new deployments, Hugging Face recommends engines such as vLLM and SGLang for servers and llama.cpp or MLX for local use. Inference Endpoints support these engines directly.
What changed in Transformers v5?
Version 5 refactored model definitions to be simpler and more modular, focused the library on PyTorch, improved built-in inference features, removed or deprecated legacy APIs, and moved to weekly minor releases. Most code using pipelines and Auto classes works with small changes; follow the official migration guide for details.
Is it safe to download models from the Hub?
Generally yes, with care. Prefer models in the safetensors format from reputable organisations, read model cards, check licenses, avoid enabling remote code you have not reviewed, and pin revisions in production. Hugging Face also runs security scans on uploaded files.
Conclusion: The Open AI Stack Is Yours to Build With
Hugging Face has grown from a chatbot app into the connective tissue of open AI. Its Hub is the library where millions of models and datasets live. Its libraries, from Transformers v5 and Datasets to PEFT, TRL, Diffusers and Sentence Transformers, are the standard tools for building and adapting models. Its inference options, from in-browser Transformers.js to serverless Inference Providers and dedicated Endpoints, cover every deployment scenario. Spaces, Gradio, smolagents and MCP take you all the way to applications and agents that people can actually use.
The most important idea to take away is that you do not need to choose between understanding and productivity. You can start with a three-line pipeline today and, as your needs grow, peel back each layer: tokenizers, chat templates, quantization, fine-tuning, serving engines. Every layer is open, documented and backed by a community that shares its work.
So pick a small, real problem from your work or studies, open the Hub, and build something this week. Publish it as a Space, write a good model card, and share it. That is how almost everyone in the open AI world got started, and in 2026 the tools have never been more powerful or more accessible.
Information reflects the state of tools, libraries and plans as of October 2026. AI products change quickly, so check official documentation and pricing pages for the latest details.

Comments
Post a Comment