Skip to content
The Injection_
LIVE
LIVElast sweep UTCnext in--:--:--

[ RSS ][ JSON ][ llms.txt ]

CAISI: GLM-5.3 is the most cyber-capable open-weight model yet

A US government eval puts an open-weight model's exploit skills well below frontier models, but ahead of other open weights.

· NIST CAISI · Research by Center for AI Standards and Innovation (CAISI) (NIST)

GLM-5.3
INFOBenchmark

NIST's Center for AI Standards and Innovation rates Z.ai's GLM-5.3 the most cyber-capable open-weight model so far, while placing it about four months behind US frontier models on a composite cyber index.

01What happened

NIST's Center for AI Standards and Innovation (CAISI) published its assessment of Z.ai's GLM-5.3 on September 17, 2026. CAISI calls it the most cyber-capable open-weight model released to date, and says it is significantly lower than US frontier models, about four months behind on a composite index.

02How it works

CAISI ran GLM-5.3 as an agent with bash and Python tools on four benchmarks: SEC-Bench Pro, ExploitBench, ExploitGym Userspace and a private OSS-Fuzz set. GLM-5.3 scored 40.4% on SEC-Bench Pro against 90.2% for a US frontier model, and 61.1% on ExploitBench against 100%. CAISI combines results into an item response theory index, where +400 points means ten times the odds of solving a task.

03Who is affected

Defenders whose threat models assume attackers lack strong exploit-development help. GLM-5.3 has open weights, so anyone can download and run it without access controls.

04What to do

Assume capable, freely downloadable exploit-writing assistance is available to attackers, and shorten patch timelines for known vulnerabilities accordingly.

Sources 2 links, all fetched

  1. 1nist.govCAISI Assessment
  2. 2anthropic.comAnthropic Analysis