AABHA AI

AI news

Kolibri adds open German–English reasoning; ThinkingBox tests agents’ actual results

Aleph Alpha’s official model card dates Kolibri-1’s release to October 3. The German–English reasoning model provides Apache 2.0 weights, tool calling and a mixture-of-experts architecture with about 78 billion total parameters and 3.46 billion active per token. The card lists a maximum context of 1,048,576 tokens but recommends no more than 262,144 for complex tasks and serving efficiency. Builders evaluating local deployment should distinguish active compute from storage: the released FP8 weights still require roughly 78 GB of model memory. For learners, it illustrates why a model can perform relatively little work per token yet remain too large for an ordinary laptop. The promising use case is specialized bilingual workflows, subject to testing on the language and documents a team actually uses.

Aleph Alpha model card and release date · Announced 3 October 2026

Microsoft and Hugging Face announced ThinkingBox availability through Hugging Face and OpenEnv. Its benchmark checks the final records and side effects left by agents across 507 business workflows, repeating each task 20 times. This distinguishes occasional success from repeatable execution and catches cases where an agent reports completion while leaving incorrect system state. The published results are the authors’ measurements, not a guarantee about production deployments. Builders can take a useful evaluation principle from the release: check the resulting database fields and unintended changes, then repeat the same scenario from a clean starting state. Learners should also distinguish pass-at-least-once metrics from success on every repeated attempt; they answer different questions about an agent’s usefulness.

Microsoft and Hugging Face announcement

All AI news