This paper introduces a new environment called Trace, which allows vision-language models to reason across multiple domains and tasks, using a taxonomy-guided approach. Practitioners might care because this work could lead to more generalizable and transferable AI models.
Firehose
Filtered to tagged “multimodal models” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News
Browse by tag
artificial intelligence 43open-weight models 26agentic coding 20AI 14reinforcement learning 14AI safety 13continual learning 13AI agents 10cybersecurity 7large language models 7open-source 6Reinforcement learning 6benchmarking 5finance 5tech 5web development 5Agentic AI 4AI ethics 4autoregressive models 4Diffusion models 4language models 4Recursive self-improvement 4security 4vision-language models 4AI infrastructure 3AI security 3Code generation 3diffusion models 3diffusion transformers 3general 3
Thinky's Inkling is a 975B-params, 41B-active-params multimodal Mixture-of-Experts transformer model, the first open-weights foundation model from the company, supporting text, images, and audio inputs with a context window of up to 1M tokens, pre-trained on 45 trillion tokens of text, images, audio, and video. Inkling is licensed under Apache 2.0 and has been released with full weights available, along with immediate support on Tinker platform and Hugging Face. The model is designed for practical use and customization, with controllable reasoning effort levels and a focus on efficient and controllable thinking. AI summary