Colibrì: Running GLM-5.2 (744B) Locally in RAM With No GPU
By Fahd Mirza
This is an AI-generated summary of “Colibrì: Running GLM-5.2 (744B) Locally in RAM With No GPU” — a 10 min YouTube video by Fahd Mirza, published July 14, 2026. It condenses the full transcript into 9 key takeaways with clickable timestamps.
Summary
This video demonstrates how to run the massive 744 billion parameter GLM 5.2 Mixture of Experts model on a single machine with limited RAM by streaming model experts from disk using a custom C engine called Colibri.
Key Points
- The video showcases running GLM 5.2, a 744 billion parameter Mixture of Experts (MoE) model, on a single machine without a cluster or multiple GPUs.
- This feat is accomplished using Colibri, a tiny C engine with zero dependencies that streams model components from disk on demand.
- The core principle leverages the fact that the model needs about 370 GB on disk but only a fraction of that (around 10 GB for dense layers plus cached experts) in RAM at any given point.
- As an MoE model, GLM 5.2 only activates a small subset of its 21,050 experts (e.g., 8 out of 256 per layer) for each token, rather than using all parameters simultaneously.
- The system's speed improves with continued use because Colibri intelligently caches frequently hit experts in RAM, leading to a higher 'expert hit rate' and faster token generation.
- The primary performance bottleneck for this system is the disk read speed, not the CPU or GPU, as the GPU remains largely idle during the generation process.
- Colibri manages memory by keeping the dense model parts and frequently accessed 'hot' experts in RAM, while streaming other necessary experts directly from disk.
- This streaming technique is specifically effective for Mixture of Experts models because only a small portion of the model is required per token, making it impractical for dense models.
- While the current Colibri implementation is handwritten and optimized for GLM 5.2's exact architecture, the underlying technique is general and holds promise for other MoE models.
Summarize any YouTube video, free
You just read an AI summary of this video. Paste any other YouTube link and get the key points with clickable timestamps in seconds — no signup, 5 free a day.
More Resources
More Summaries
23 minPoE 3.29 - Ice Crash Ignite Chieftain - Build Guide
This video details an "Ice Crash Ignite Chieftain" build for Path of Exile's 3.29 league, highlighting its overpowered status, insane clear speed, strong single-target damage, and robust defenses as a
4 minIndian Visa Appointment Booking Online | Step-by-Step IVACBD Portal Guide
This video provides a step-by-step guide on how to book an Indian visa appointment online through the IVAC BD portal, emphasizing accurate data entry and timely actions.
2 minHoly Spirit Fight for Me #inspiration #motivation #love
This video is a fervent prayer invoking the Holy Spirit to fight spiritual battles across all aspects of life, declaring victory and rejecting defeat through divine intervention.
55 minClaude Code built me a $273/Day online directory
This video provides a comprehensive guide on building profitable online directories with minimal investment and effort, leveraging AI tools like Claude Code and Crawl for AI to automate data acquisiti
6 minGSP teaches Lex Fridman how to street fight
Georges St-Pierre shares essential self-defense tactics for street fights, emphasizing the critical role of surprise, striking vulnerable points, and strategic responses to various threats, including