FlashLabs Introduces MLX Quantized LLM for Apple Silicon
FlashLabs, a prominent AI research company based in Chiyoda, Tokyo, has made headlines with the release of a new MLX quantized version of the open-weight large language model (LLM) known as GLM-5.3-Flash by Z.ai. Optimized especially for Apple Silicon, this model is available in five different quantized variants ranging from "2-bit Lite" to "6-bit." With a total parameter count of 320 billion, and 18 billion active parameters, this model is designed to make advanced AI functionalities more accessible to developers and companies.
Key Features
The GLM-5.3-Flash-MLX package boasts several impressive characteristics:
- - Massive Parameter Sizes: It has a total of 320 billion parameters, which includes 18 billion active parameters for improved performance.
- - Multiple Variants: Consumers can choose from five variants: "2-bit Lite," "2-bit," "3-bit," "4-bit," and "6-bit." Each variant presents different compression levels, catering to various memory capacities and performance needs.
- - Unique Quantization Method: Utilizing a proprietary quantization technique called "OrcaSAQ," the model ensures accuracy without requiring calibration data. This method preserves critical tensors based on the model's architecture, achieving a balance between performance and memory usage.
- - Minimal Sacrifice on Performance: The highest tier, the "6-bit" version, claims to be nearly lossless when compared to its FP8 counterpart, with a Top-1 accuracy rate of 97.76%.
- - Lightweight Option Available: The "2-bit Lite" variant, specifically refined for lower memory systems, requires just 112GB of memory and can function on the MacBook Pro featuring 128GB of memory, making advanced AI applications feasible for users.
Development Journey
The local execution of large language models has gained traction recently, helping developers manage data more securely without the need for external transmission and reducing latency. Despite this appeal, many existing models tend to demand hundreds of gigabytes of memory, making them impractical for everyday laptops.
In response to these challenges, FlashLabs has focused on compressing model sizes while maintaining performance through innovative quantization technologies. The introduction of OrcaSAQ is a hallmark of this strategy, showcasing the company’s commitment to refining AI capabilities while ensuring ease of access for users.
Variants Breakdown
Following the introduction of the five variants on August 27, 2026, the specifications include:
| Variant | Model Size | Required Memory | Key Features |
|---|
| -- | -- | ---- | --- |
| 6-bit | ~296GB | 320GB | Nearly lossless (Top-1 accuracy 97.76%) |
| 4-bit | ~204GB | 224GB | Good performance; recommended default |
| 3-bit | ~184GB | 200GB | Good performance (aggressive) |
| 2-bit | ~145GB | 160GB | Aggressive; best-effort option |
| 2-bit Lite | ~102GB | 112GB | Minimal; compatible with MacBook Pro |
The lightest version, "2-bit Lite," significantly compresses the model size to about 102GB, fitting within the limits of a MacBook Pro equipped with the M4/M5 Max chips. This variant retains critical functionalities of the GLM-5.3-Flash model, including the capacity for 1 million tokens context and multimodal input. It is also capable of running on a single GPU (NVIDIA H200) with sufficient memory for operations.
Public Access and Future Plans
The GLM-5.3-Flash-MLX is freely available under the MIT license on Hugging Face, allowing anyone to download and utilize the model. For those in environments where local execution presents challenges, the model can be accessed with full precision through the OrcaRouter API.
Looking ahead, OrcaRouter is dedicated to continuing research and development in quantization technologies to empower users to harness large models on everyday devices. The aim is to expand the model catalog further and provide users with versatile AI toolkits to enhance their operations.
About OrcaRouter
OrcaRouter is an AI inference gateway developed by Continuum AI, a U.S.-based research organization, with exclusive distribution in the Japanese market by FlashLabs. Equipped with OpenAI-compatible APIs, it allows users to access over 200 LLMs and generative AI models from a single entry point. Recent additions to the catalog include the latest models from prominent developers, ensuring users have access to cutting-edge solutions in AI.
Company Overview
- Headquarters: Chiyoda, Tokyo, Japan
- CEO: Yoichi Hosoi
- Business Focus: AI application research
- Website:
FlashLabs
- Mission: Developing foundational infrastructure for intelligent systems for the next decade
- Website:
Continuum AI
For press inquiries, please contact: Koki Kobayashi, Marketing Department at FlashLabs via email at
[email protected].