The Hugging Face team has released a new version of the tokenizers library (v1) with improved performance, including a hand-written splitter instead of a regex engine, a cache that answers repeated words without merging, and a merge loop that never touches the allocator. The library now scales across cores and processes multiple pre-token spans in one call. The new version is faster than the previous version (v0.23) by 3-30 times, depending on the model and input size. AI summary
Firehose
Filtered to tagged “tokenizers” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
artificial intelligence 82continual learning 29AI 26agentic coding 17reinforcement learning 16open-weight models 11AI agents 10AI safety 9AI ethics 8cybersecurity 8machine learning 7open-source 7language models 6natural language processing 6Reinforcement learning 6artificial general intelligence 5deep learning 5Diffusion models 5Agentic AI 4computer vision 4conversational AI 4ethics 4existential risk 4large language models 4LLMs 4Recursive self-improvement 4robotics 4security 4vision-language models 4ai 3