AI & ML interests

The AI community building the future.

Recent Activity

nielsrย  updated a bucket about 3 hours ago
huggingface/paperswithcode-backups
tarekziadeย  updated a bucket about 6 hours ago
huggingface/transformers-ci-telemetry
alvarobarttย  updated a dataset about 6 hours ago
huggingface/DEH-image-scan-data
View all activity

Articles

nielsrย 
updated a bucket about 3 hours ago
tarekziadeย 
updated a bucket about 6 hours ago
sergiopaniegoย 
posted an update 1 day ago
view post
Post
93
catching up on some bookmarked reads from the summer, reading Antidoom from @liquidai

small reasoning models get stuck more easily when the task involves a long thinking trace and a hard problem. It starts repeating the same word over and over again ("Wait", "Alternatively"โ€ฆ), each repetition makes the next one likelier, and the generation is spent before it reaches an answer

they measured it, 10.2% of completions for an early LFM2.5-2.6B checkpoint and 22.9% for Qwen3.5-4B at greedy. After training those drop to 1.4% and 1.0%

the fix is FTPO (final token preference optimization). What I like is how narrow it is, it only touches the single token where the loop starts

three ways it differs from DPO:
> trains one token position, mid-generation, instead of whole sequences
> spreads probability across ~20 plausible alternatives instead of swapping one overtrained token for another
> keeps the regularizer in logit space, no softmax, so the rest of the vocabulary stays put

the third one is what makes it usable. If you want to edit one position without disturbing the model, you can't have a loss that reshuffles the other 150k logits on the way

and their explanation abt the result: the training teaches the model nothing new about math or code, it clears the failure mode that was blocking answers the model could already produce

full blog > https://www.liquid.ai/blog/antidoom

FTPO itself comes from Antislop, where it was built to strip overused phrasing. LiquidAI retargeted it to doom loops

and under the hood it's a subclass of TRL's DPOTrainer with compute_loss overridden, around 90 lines of loss and no new trainer

we documented that pattern in TRL's docs
https://huggingface.co/docs/trl/main/en/customization#change-the-training-objective
sergiopaniegoย 
posted an update 10 days ago
view post
Post
210
super interesting new paper from Microsoft "Agent Lightning v1.0: Towards Harnessed Agentic RL" by Zhiyuan He et al.

same idea we've seen already several times: you train the agent inside the real harness it ships with, instead of a reimplementation of it

now that recipe has a name โ†’ harnessed agentic RL

paper: huggingface.co/papers/2608.17528

the tricky bit they nail down: one rollout is not one training sample

the harness calls the model many times, so a single episode โ†’ a variable number of (prompt, response) rows

you don't even know the batch size until the episode finishes running

its real contribution is being first to systematically map the four problems that fall out of that:

> retokenization + sample merging
> advantage calculation over a variable sample count
> loss normalization at the rollout level, not per sample
> backend scheduling when the batch size is dynamic

and it actually works โ†’ plain RL inside the real harness, no reimplementation

Qwen3.5-9B on SWE-bench Verified 41.8 โ†’ 56.4 (+14.6), with only ~6k examples

the whole thing is ~3,500 lines, any harness, self-hosted k8s

from our side, we've shared some materials on the same line you may want to check out :)

> Agentic RL: Token-In, Token-Out Done Right: https://huggingface.co/blog/huggingface/tito
> a full worked example, opencode owning its loop trained with GRPO: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
> Harness, Scaffold, and the AI Agent Terms Worth Getting Right: https://huggingface.co/blog/agent-glossary

on a similar line:

https://x.com/SergioPaniego/status/2062911580564496576
tomaarsenย 
posted an update 12 days ago
view post
Post
3686
๐Ÿšจ I've just published Sentence Transformers v6.0, introducing MultiVectorEncoder: ColBERT-style late interaction models are now a fourth model type, for training, inference, and interpretation, alongside the dense, sparse, and reranker models! Details:

Where a regular embedding model compresses a whole text into one vector, a multi-vector model keeps one vector per token and scores query against document with the MaxSim operator. That preserves token-level matching information that a single vector has to average away. It is also the state of the art for visual document retrieval, where a text query is matched against page images directly, charts and tables included, with no OCR step in between.

Any PyLate, Stanford ColBERT, or ColPali checkpoint loads straight into the same familiar API: model.encode_query(), model.encode_document(), and model.similarity() just work, whether the documents are texts or page images.

Does it help? LightOn trained LateOn (multi-vector) and DenseOn (dense) on the same data with the same 149M ModernBERT backbone, and the multi-vector model wins on 9 of the 13 NanoBEIR datasets: 0.6868 vs 0.6764 mean NDCG@10. The price is a bigger index, and the new HierarchicalTokenPooling module halves it at roughly no retrieval cost.

Antoine Chaffin, Raphaรซl Sourty, and I wrote a blog post walking through multi-vector models in practice: loading the various checkpoint formats, encoding and scoring, plugging them into a search stack, running them on page images, and keeping the index affordable. Check it out if you want to get started, or just point your Agent to the URL: https://huggingface.co/blog/multi-vector-encoder

pip install sentence-transformers==6.0.0

Release notes: https://github.com/huggingface/sentence-transformers/releases/tag/v6.0.0
sergiopaniegoย 
posted an update 20 days ago
view post
Post
695
Something I really like when I study a subject is understanding its history, how it reached the point where it is today

I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words

This is the companion piece to Class 3 of our Training Agents series with @burtenshaw . The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale

https://huggingface.co/blog/sergiopaniego/agentic-rl-2026
  • 1 reply
ยท
sergiopaniegoย 
posted an update 25 days ago
view post
Post
288
we just released a new blog "Training a coding agent using the OpenCode harness in remote HF sandboxes with TRL and OpenEnv"

you can take a real coding agent (OpenCode), let it run its own tool loop against real coding problems, and train it with RL on the exact tokens it produced

and every rollout runs in its own remote HF sandbox, so rollouts scale out beyond one machine

the loop:
- OpenCode owns its tool loop inside an OpenEnv sandbox
- an in-sandbox proxy records the real token ids + logprobs, per turn
- a hidden-test verifier scores the result, and that is the reward
- TRL trains with AsyncGRPO, weights sync back to vLLM over NCCL

blog + runnable example: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
  • 3 replies
ยท
sergiopaniegoย 
posted an update 26 days ago
view post
Post
2620
LFM2.5-2.6B just dropped!

and the @liquidai blog comes with some nice details about the training procedure, so let's analyze it.

basically, a full agent training pipeline but compressed into 2.6B

base model โ†’ SFT โ†’ specialized teachers per domain (SFT + RLVR) โ†’ on-policy distillation back into one student โ†’ agentic RL

the two most interesting stages

โ†’ MOPD: the student generates, each prompt routes to its domain teacher for token-level feedback. teachers branch from the same SFT checkpoint, so their signal stays close to the student's distribution

โ†’ agentic RL: multi-turn GRPO inside real harnesses (OpenClaw, Hermes Agent), one sandbox per rollout, a proxy captures token-level trajectories while the harness stays a black box

this makes a 2.6B that beats much larger models on instruction following and tool use

SFT, distillation, RL, RL envs: exactly what we're covering in our Training Agents livestream series (next one coming soon!)

โ†’ model: LiquidAI/LFM2.5-2.6B
โ†’ blog: https://www.liquid.ai/blog/lfm2-5-2-6b
โ†’ live series: https://www.youtube.com/playlist?list=PLo2EIpI_JMQvQZm-kVlz4wY1vWF0LBcf5
  • 3 replies
ยท
sergiopaniegoย 
posted an update about 1 month ago
view post
Post
2652
Simon Willison (@simonw ) has asked every new model to draw a pelican riding a bicycle for some time now

you look at the drawing and you know. but there is no number, so nothing can train against it, no?

I turned this idea into an rl env in OpenEnv. now, you can eval any model against it, and train against it with TRL

read the details!๐Ÿค“

https://huggingface.co/blog/sergiopaniego/pelican-env-openenv
  • 2 replies
ยท
sergiopaniegoย 
posted an update about 1 month ago
julien-cย 
posted an update about 1 month ago
view post
Post
4659
who's working on an NVFP4 version of Kimi-K3?
  • 4 replies
ยท
sergiopaniegoย 
posted an update about 1 month ago
view post
Post
2917
quick reminder! ๐Ÿšจ

tomorrow (Tuesday, July 28), we're back with Class 3 of the Training Agents live series

๐Ÿง  what: reinforcement learning for training agents (GRPO): how it works, how to implement it in TRL, and end-to-end examples
๐Ÿ—“๏ธ when: Tuesday, July 28 - ๐Ÿ•” 5:00 PM CEST / 8:30 PM IST
๐Ÿ“ where: Live on @huggingface 's X, YouTube, and LinkedIn

live: https://www.youtube.com/watch?v=ztdTed5egrM

class 1: https://x.com/SergioPaniego/status/2069382207618379813
class 2: https://x.com/SergioPaniego/status/2075180665184686187
  • 1 reply
ยท
sergiopaniegoย 
posted an update about 1 month ago
view post
Post
238
you can now train your own coding agents with trl + openenv, starting with opencode

we just added end-to-end support for training agent harnesses:

> TRL: a loop-owning training path (AsyncGRPOTrainer + HarnessRolloutWorker) that launches the agent in an OpenEnv session, reads back its trace, reconstructs the training samples, and trains with AsyncGRPO
> OpenEnv: the OpenCode harness environment plus a transparent proxy that forwards the agent's model calls and records each turn's token ids and logprobs

you train the actual opencode agent as is, it runs its own loop and tools and the policy learns from the exact tokens it produced

we're shipping a self-contained example: local subprocess sandbox, DeepCoder problems, validated on Qwen3-8B.

> example: https://github.com/huggingface/trl/blob/main/examples/scripts/openenv/opencode.py
> docs: https://huggingface.co/docs/trl/main/openenv

and we're working actively on both sides so expect more ๐Ÿค“
  • 1 reply
ยท
badaouiย 
posted an update about 1 month ago
view post
Post
2388
432 GB of ultra-fast HBM4 and up to 23.3 TB/s of memory bandwidth on a single GPU ๐Ÿคฏ.

Two weeks ago, we got early access to AMD's new Instinct MI455X, and our first goal was simple: make sure ๐Ÿค— Transformers works on day one.

Over the past few weeks, we worked closely with the AMD team to validate the platform, enable Flash Attention, add torchcodec support for multimodal models, and resolve issues uncovered during testing.

The result:
โœ… 99.5% success rate across our 24 core Transformers model architectures - already on par with our daily CI on previous AMD and NVIDIA platforms.

The hardware is just as exciting. With 432 GB of HBM per GPU, our early capacity experiments showed more than 3ร— the concurrent long-context requests compared to MI300, thanks to the much larger KV cache capacity.

A huge thanks to the AMD team for the early access and the great collaboration!

Read the full blog ๐Ÿ‘‡
https://huggingface.co/blog/badaoui/transformers-on-amd-mi455
  • 1 reply
ยท