BREAKING!
Mumbai’s Vile Parle Cypher is giving unheard rappers a voice Enrich and Ultimo impress Karnataka CM Signals Review of 5% Park Land Amendment Bill in Assembly When an organisation deviates from its values, faith of those associated with its ideology is shattered: Kaushal Kishore Andhra Pradesh CM Urges Farmers to Cede Land for Narsapuram Port Development Rajasthan local body polls: Congress launches ‘Gen Z for Change’ campaign At least 30 houses set ablaze in three Manipur villages; 1 injured: Police CBI Faces 7,229 Pending Corruption Cases, 409 Over 20 Years Old ‘Cauvery our legal right’: Tamil Nadu CM Vijay opens Mettur dam for Samba rice cultivation, drinking water needs MP woman, 21, murdered by father, uncle for working outside, short hair: Police Mumbai’s Vile Parle Cypher is giving unheard rappers a voice Enrich and Ultimo impress Karnataka CM Signals Review of 5% Park Land Amendment Bill in Assembly When an organisation deviates from its values, faith of those associated with its ideology is shattered: Kaushal Kishore Andhra Pradesh CM Urges Farmers to Cede Land for Narsapuram Port Development Rajasthan local body polls: Congress launches ‘Gen Z for Change’ campaign At least 30 houses set ablaze in three Manipur villages; 1 injured: Police CBI Faces 7,229 Pending Corruption Cases, 409 Over 20 Years Old ‘Cauvery our legal right’: Tamil Nadu CM Vijay opens Mettur dam for Samba rice cultivation, drinking water needs MP woman, 21, murdered by father, uncle for working outside, short hair: Police
LEADNEWS24

LM Studio Adds GLM-5.3-Flash to Bionic with Image Support & 1M-Token Context

LM Studio integrates Z.ai’s GLM-5.3-Flash into Bionic, offering multimodal input, 1M-token context, and 10x lower costs than GLM-5.2. Available now on Mac and Windows.

9
9to5Mac
Aug 28, 2026 · 4 min read
LM Studio Adds GLM-5.3-Flash to Bionic with Image Support & 1M-Token Context

LM Studio announced that its AI agent platform, Bionic, now supports Z.ai’s GLM‑5.3‑Flash model, adding multimodal input, a one‑million‑token context window and a price that is up to ten times lower than its predecessor, GLM‑5.2. The update expands Bionic’s already growing catalog of models, which can run locally on a user’s Mac or Windows machine or be accessed through cloud servers located in the United States under a strict zero‑data‑retention policy.

The GLM‑5.3‑Flash model, also known as “Ox Alpha,” is a 320‑billion‑parameter mixture‑of‑experts architecture with 18 billion active parameters. It can process both text and image inputs and supports a context window of up to one million tokens, a significant increase over the 256‑k token limit of earlier models. According to Z.ai’s benchmark results, GLM‑5.3‑Flash outperforms GLM‑5.2 on a range of tests and sits competitively with frontier models from Anthropic, OpenAI, Google and DeepSeek.

LM Studio’s tweet on August 26, 2026, announced the new cloud support: “GLM‑5.3‑Flash by @Zai_org, aka Ox Alpha, is now live in LM Studio Bionic! This model surpasses GLM 5.2 in performance while being 9‑10x cheaper. It also supports image input. In Bionic, this model is served from US‑based servers with ZDR enabled by default. Happy building!” The company noted that the model can be accessed through Bionic’s cloud interface, which automatically enables ZDR (Zero‑Data‑Retention) for added privacy.

The update follows a series of model additions that began with the launch of Bionic in mid‑July. Within weeks, LM Studio added support for Moonshot AI’s Kimi K3 model, and more recently the company expanded its cloud offerings to include Z.ai’s GLM‑5.3‑Flash. The platform’s hybrid architecture allows developers and researchers to choose between running large models on local hardware—if they have the necessary GPU resources—or offloading inference to LM Studio’s US‑based servers.

GLM‑5.3‑Flash was officially unveiled by Z.ai earlier in the week, sparking interest from the developer community. The model has been independently tested on OpenCode and OpenRouter under the codename “Ox Alpha,” where it demonstrated higher accuracy and faster response times compared to GLM‑5.2. While the exact pricing tiers are not disclosed in the announcement, LM Studio’s statement that the model is “up to 10 times cheaper” suggests a significant cost advantage for users who previously relied on GLM‑5.2 or other large‑language‑model services.

Bionic’s new multimodal capabilities enable a range of use cases that extend beyond traditional text‑based tasks. Developers can now build agents that interpret images alongside textual prompts, making it possible to automate workflows that involve document analysis, visual data extraction, or image‑guided coding. The one‑million‑token context window also allows Bionic to maintain longer conversations or process larger documents without truncation, improving the continuity and relevance of generated responses.

The announcement comes at a time when the demand for affordable, high‑performance AI models is growing across industries. LM Studio’s focus on privacy—through its zero‑data‑retention policy for cloud models—aligns with regulatory trends that prioritize data protection. By offering a cloud‑hosted version of GLM‑5.3‑Flash that is both cheaper and privacy‑preserving, the company positions itself as a viable alternative to larger, more expensive providers.

Users who run local models on their Macs are encouraged to share their experiences in the comments. For those interested in integrating GLM‑5.3‑Flash into Bionic, LM Studio provides a link to documentation and setup instructions. The platform’s continued expansion suggests that it will remain a key player in the evolving landscape of AI agent development.

Source: 9to5Mac. Rewritten by AI · How We Use AI →
View original source →

Comments (0)

Comments are moderated and may take a little while to appear.

No comments yet — be the first to weigh in.

Related
Trending

We use cookies to improve your experience and analyze traffic.

Manage cookie preferences

Essential

Required for the site to function. Always active.

Analytics

Helps us understand how readers use the site.

Marketing

Used to personalize ads shown to you.