English

Model ReleasesHugging Facetokenizers

Hugging Face Releases tokenizers v1 Release Candidate with Significant Performance Gains

Hugging Face has released the release candidate for tokenizers v1, an update focused on significantly improving performance to prevent tokenization from becoming a bottleneck in machine learning workflows. The new version maintains full compatibility with v0.23, producing the same token IDs, API, and vocabulary.

The performance improvements stem from several key architectural changes. The library now uses a hand-written splitter instead of a general-purpose regular expression engine, utilizing SIMD (single instruction, multiple data) instructions to process multiple bytes at once. Additionally, v1 implements a thread-local cache that stores results for previously processed pre-tokens, allowing it to skip the expensive merge process for repeated words. The merge loop has also been optimized to reuse scratch buffers, eliminating frequent memory allocations, and it now processes batches of pre-tokens in a single model call.

Benchmarks conducted on an Apple M4 Max show that the v1 encode path is 3 to 30 times faster than v0.23 in single-threaded mode, depending on the model family. The update also scales effectively across multiple cores, achieving approximately 76% of linear scaling across eight workers.

The release candidate is available on crates.io. Future updates are expected to expand support for additional model families and eventually bring these improvements to the transformers library.

Sources

  1. tokenizers v1: encode, decode and scaling, measured (Hugging Face Blog, 2026-09-21)
  2. Update on GitHub