Colibri: Running a 744B Parameter Model on 25GB of RAM

Discover Colibri: a tiny C engine that streams experts from disk to run massive 744B parameter models like GLM-5.2 on machines with only 25GB of RAM.

Discover Colibri: a tiny C engine that streams experts from disk to run massive 744B parameter models like GLM-5.2 on machines with only 25GB of RAM.

Generate holistic, privacy-first documentation for large codebases locally with CodeWiki's Python CLI. Supports 9 languages and Mermaid diagrams.

Stop wrestling with generic prompts. Install a complete team of specialized AI experts into Claude Code, Cursor, and more instantly.

A week-long experiment with 100+ agents optimizing Gemma 4 revealed surprising social emergence, self-policing, and a 5x speed improvement in vLLM.

Move beyond simple LLM prompts. Build secure, autonomous agents with Flue's programmable TypeScript harness and integrated sandboxing.

Stop your LLM agents from forgetting tools. Learn how heku uses lazy discovery to manage hundreds of MCP servers without exhausting the context window.

Master your data with RAGFlow, an open-source engine that combines deep document understanding with agentic workflows to eliminate AI hallucinations.

Solve the context gap in LLM agentic systems with OKF, a vendor-neutral standard for portable, human- and machine-readable knowledge.