Research
Benchmark results, architecture deep-dives, and memory engineering insights.
August 23, 202611 min
75.5% on BEAM-100K: What Reasoning Effort Does to a Memory Pipeline
A stealth reader and one API parameter took our BEAM-100K average score from 69.3% to 75.5% (323/400 correct). Not SOTA. The interesting part is where the points came from, and where they refused to come from.
BEAMOx Alphareasoning effortColBERT
August 3, 20268 min
Memory Quality Beats Reader Quality: Evidence from BEAM-100K
How reader-agnostic memory engineering closed a 22-point gap on BEAM with GLM-5.2, an open-weights model that scores higher on agentic intelligence benchmarks than Gemini 3.1 Pro.
BEAMbenchmarkGLM-5.2filtered-diskann

