# CAISI: GLM-5.3 is the most cyber-capable open-weight model yet

> A US government eval puts an open-weight model's exploit skills well below frontier models, but ahead of other open weights.

- **Severity:** Info: Research, tools, evals or lab claims with no direct exposure. Context, not an alert.
- **Category:** Benchmark
- **Published by:** NIST CAISI
- **Disclosed:** 2026-09-17
- **Affects:** GLM-5.3
- **Research by:** Center for AI Standards and Innovation (CAISI) (NIST)
- **Primary source:** https://www.nist.gov/news-events/news/2026/09/caisis-assessment-zais-glm-53-cyber-capabilities
- **Page:** https://theinjection.dev/items/caisi-glm-5-3-cyber-assessment/

NIST's Center for AI Standards and Innovation rates Z.ai's GLM-5.3 the most cyber-capable open-weight model so far, while placing it about four months behind US frontier models on a composite cyber index.

## What happened

NIST's Center for AI Standards and Innovation (CAISI) published its assessment of Z.ai's GLM-5.3 on September 17, 2026. CAISI calls it the most cyber-capable open-weight model released to date, and says it is significantly lower than US frontier models, about four months behind on a composite index.

## How it works

CAISI ran GLM-5.3 as an agent with bash and Python tools on four benchmarks: SEC-Bench Pro, ExploitBench, ExploitGym Userspace and a private OSS-Fuzz set. GLM-5.3 scored 40.4% on SEC-Bench Pro against 90.2% for a US frontier model, and 61.1% on ExploitBench against 100%. CAISI combines results into an item response theory index, where +400 points means ten times the odds of solving a task.

## Who is affected

Defenders whose threat models assume attackers lack strong exploit-development help. GLM-5.3 has open weights, so anyone can download and run it without access controls.

## What to do

Assume capable, freely downloadable exploit-writing assistance is available to attackers, and shorten patch timelines for known vulnerabilities accordingly.

## Quick facts

- **SEC-Bench Pro:** 40.4% vs 90.2% (US frontier)
- **ExploitBench:** 61.1% vs 100% (US frontier)
- **ExploitGym Userspace:** 9.4% (47 of 502)
- **CAISI OSS-Fuzz:** 7.7% (23 of 297)
- **Gap to frontier:** About four months (composite index)

## Sources

1. [CAISI Assessment](https://www.nist.gov/news-events/news/2026/09/caisis-assessment-zais-glm-53-cyber-capabilities)
2. [Anthropic Analysis](https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities)
