GLM-5.3-Flash Uncensored
2026-09-02 02:50:38

FlashLabs Unveils QC Version of GLM-5.3-Flash-Uncensored for OrcaRouter

FlashLabs Introduces New GGUF Model



FlashLabs, located in Chiyoda, Tokyo, has announced a significant development in the AI domain. The company has unveiled the GGUF quantized version of the model "GLM-5.3-Flash-Uncensored" aimed at security researchers to enhance both Red Team and Blue Team capabilities. This model is part of OrcaRouter, an AI inference gateway that FlashLabs exclusively offers in the Japanese market, developed by the US-based AI research organization Continuum AI.

Overview of the New Model



The GLM-5.3-Flash-Uncensored model builds on the open-weight large language model (LLM) called GLM-5.3-Flash, announced by Z.ai, which boasts a massive 320 billion parameters and 18 billion active ones. The newly released GGUF version specifically targets research applications, having relaxed content filtering to facilitate better usability for security professionals. Following the release of the FP8 version back in late August 2026, this new GGUF variant comes in four types: Q6_K, Q4_K_M, Q3_K_M, and Q2_K, with Q4_K_M recommended as the most balanced option.

Key Features



The model variants possess different file sizes ranging from 263GB for Q6_K to 117GB for Q2_K, ensuring flexibility based on user requirements. Furthermore, they support both CPU and GPU environments, including compatibility with Apple Silicon, enabling them to run locally without relying on cloud services for sensitive data.

In addition, the GGUF version integrates a multimodal visual encoder (mmproj) that extends its capability to handle image input, enhancing its versatility in security validations. A significant advantage of this model is the incorporation of weight adjustments directly into the model, avoiding dependencies on jailbreak prompts or supplementary learning for response rejection.

Variant Specifications


Variant File Size Top-1 Accuracy
-----------
Q6_K 263GB 94.3%
Q4_K_M 193GB 89.8% (recommended)
Q3_K_M 153GB 85.6%
Q2_K 117GB 75.5%

All four variants can handle both CPU and GPU environments, making them adaptable for diverse hardware configurations, including local environments for security operations.

Responsible Usage



Continuum AI emphasizes the importance of responsible utilization of the GLM-5.3-Flash-Uncensored model, limiting its application to verification and research purposes by security researchers. To access this model, users must submit a request and adhere to licensing terms available on Hugging Face. Furthermore, when utilized in production environments, safeguards, such as PII shielding, content policy enforcement, and agent firewalls, are automatically enabled, delineating research from operational use.

Future Outlook



Moving forward, OrcaRouter aims to solidify its contributions to the security research community while fostering a secure environment for AI utilization. As the landscape of AI continues to evolve, FlashLabs is committed to equipping researchers with tools that enhance both safety and practicality in their work.

About OrcaRouter



OrcaRouter is developed by Continuum AI and exclusively offered in Japan through FlashLabs. This robust AI inference gateway enables access to over 200 LLMs and generative AI models, including the latest releases from acclaimed companies such as OpenAI and Google. Users can seamlessly manage various models from a single API, simplifying AI integrations for diverse applications.

Company Details


  • - FlashLabs, Inc.
Headquarters: 10-8 Ichiban-cho, Chiyoda, Tokyo
CEO: Hiroichi Hosoi
Established: July 2023
Website: FlashLabs

  • - Continuum AI



画像1

Topics Other)

【About Using Articles】

You can freely use the title and article content by linking to the page where the article is posted.
※ Images cannot be used.

【About Links】

Links are free to use.