Newsletter image

Subscribe to the Newsletter

Join 10k+ people to get notified about new posts, news and tips.

Do not worry we don't spam!

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

Search

GDPR Compliance

We use cookies to ensure you get the best experience on our website. By continuing to use our site, you accept our use of cookies, Privacy Policy, and Terms of Service.

AI2 OLMo 3.1

The Allen Institute for AI (AI2) has released OLMo 3.1, a family of fully open 32 billion parameter reasoning models. Unlike most open-weight releases, OLMo 3.1 ships with the complete training data, training code, evaluation scripts, and intermediate checkpoints, making it the most transparent frontier model ever published. The Think 32B variant gains 5+ points on AIME, 4+ on ZebraLogic, 4+ on IFEval, and over 20 points on IFBench compared to OLMo 3. AI2 calls the Instruct 32B variant the most capable fully open chat model to date.

Alibaba Qwen3.5

Alibaba has released Qwen3.5, headlined by a 397 billion parameter Mixture-of-Experts model with 17 billion active parameters per token. Shipped under Apache 2.0 on Hugging Face, Qwen3.5 scores 93.3% on AIME 2026, 85.0 on LiveCodeBench v6, and 76.8% on SWE-Bench Verified, putting it in frontier territory for math, coding, and agent tasks. The broader Qwen3.5 family spans dense models from sub-1 billion up to 32 billion parameters, plus sparse MoE variants, giving developers open-weight options at every scale.

Moonshot Kimi K2.5

Moonshot AI has released Kimi K2.5, a 1 trillion parameter Mixture-of-Experts language model with 32 billion active parameters per request. The model uses 61 layers with 384 experts and sparse 8-expert activation, natively trained on roughly 15 trillion mixed vision and text tokens for a 256K context window. Kimi K2.5 outperforms GPT-5.2 on MMMU Pro (78.5%), BrowseComp (74.9%), and AIME 2025 (96.1%), and the Agent Swarm configuration reaches 50.2% on Humanity Last Exam at 76% lower cost than Claude Opus 4.5. Weights ship on Hugging Face under a Modified MIT license.

Z.ai GLM-5.1

Z.ai (formerly Zhipu AI) has released GLM-5.1, an open-weight 744 billion parameter Mixture-of-Experts model with 40 billion active parameters. The model immediately took the top spot on SWE-Bench Pro with 58.4, beating GPT-5.4 (57.7) and Claude Opus 4.6 (57.3). Built on DeepSeek Sparse Attention with a 200K context window and 131K maximum output, GLM-5.1 is engineered for sustained autonomous coding agents capable of running plan, execute, test, fix, optimize loops for up to eight hours. The weights are available on Hugging Face under a permissive MIT license for full commercial use.

Google Gemma 4

Google DeepMind has released Gemma 4, a family of open-weight language models in four sizes (E2B, E4B, 26B MoE, and 31B Dense) under the Apache 2.0 license. The new release brings dramatic benchmark gains over Gemma 3, with AIME 2026 math jumping from 20.8% to 89.2%, LiveCodeBench coding from 29.1% to 80.0%, and GPQA science from 42.4% to 84.3%. The flagship 31B Instruct variant ranks #3 on Arena AI text leaderboard at 1452 Elo, outperforming closed models twenty times its size. Gemma 4 ships with day-one support for Hugging Face, Kaggle, Ollama, and Google Cloud Vertex AI.

Cogito v1

DeepCogito, based in San Francisco, has introduced the Cogito v1 Preview series of open-source AI models, available in various parameter sizes from 3B to 70B. These models are designed for diverse tasks, from lightweight to heavy-duty challenges, and are freely accessible for commercial use on platforms like Hugging Face and Ollama. They claim to outperform competitors like LLaMA and Qwen, though specific benchmark scores are not disclosed. The models feature hybrid reasoning modes, enabling them to switch between standard and reasoning tasks, and are optimized for coding, STEM tasks, and agentic applications. Trained on over 30 languages with a 128k context length, they are versatile and globally applicable. While these models are early versions, larger ones are anticipated. DeepCogito encourages innovators to explore and utilize these models to unlock AI's potential and drive future advancements.

DeepSeek V3-0324

DeepSeek V3-0324, launched on March 24, 2025, is a 671 billion-parameter AI model from DeepSeek, notable for its Mixture-of-Experts architecture that activates only 37 billion parameters per token, enhancing efficiency. Competing with models like GPT-4o and Claude 3.5 Sonnet, it excels in reasoning, code generation, and multilingual tasks. Despite its size, it can run on high-end consumer PCs using optimizations like 4-bit quantization, requiring components like NVIDIA RTX 4090 GPUs, 64-128 GB RAM, and fast NVMe SSDs. Although running this model on consumer hardware involves trade-offs in speed and complexity, it remains accessible thanks to its open-source MIT license, offering a democratizing force in AI development. Future updates may enhance efficiency for consumer setups, leveraging community insights and potential model refinements.

OLMo 2 32B

The Allen Institute for AI has launched OLMo 2 32B, a 32-billion-parameter model that surpasses GPT-3.5 and GPT-4o mini on key benchmarks. Released on March 13, 2025, it is fully open-source, offering model weights, training code, datasets, logs, checkpoints, and evaluation tools. Trained on Google's Augusta hypercomputer, its performance is impressive, especially in reasoning, math, and challenge benchmarks. OLMo 2 32B promotes open AI, allowing researchers, startups, and hobbyists to explore and build upon it. Despite risks of misuse, its open-source nature holds potential for educational tools and scientific advancements. The model is available on Hugging Face, encouraging community involvement and innovation.

Gemma 3: Google Open-Source Gambit

Google's new AI model, Gemma 3, aims to outperform rivals while operating on a single GPU, making AI more accessible to developers and startups. Built on Google's Gemini 2.0 architecture, it promises high performance and versatility, handling text, images, and videos efficiently. However, despite being labeled "open," its use is restricted by an Apache 2.0 license with limitations on commercial use and redistribution. This has sparked debate in the open-source community, with some seeing it as a controlled offering rather than true open-source freedom. While early adopters have found innovative uses for Gemma 3, its constraints pose challenges for larger-scale applications. The broader question remains whether tech giants like Google can fully embrace open-source principles without relinquishing control.

Claude 3.7 Sonnet

Anthropic's Claude 3.7 Sonnet, launched on February 24, 2025, is the first hybrid reasoning AI model, blending fast responses with step-by-step thinking, similar to human cognitive flexibility. It excels in tasks like coding, math, and data analysis, achieving 70.3% on the SWE-bench Verified in standard mode. The model has two modes: standard for quick replies and extended thinking for in-depth analysis, with the latter available only on paid plans. It's useful in software development, education, and customer service, with a tool called Claude Code aiding developers. While it faces accessibility debates due to its paid features, it sets a precedent for integrating reasoning capabilities within a single AI system, with potential future enhancements and integration with other AI technologies.

Mistral AI

Mistral's chatbot offers scalable solutions for businesses, developers, and educators. Transform your operations today!