Reflection has unveiled Beam, with 501 billion total parameters and 23 billion active parameters. The company says its advanced reasoning scores approach those of GLM-5.2 while using just one-third to one-quarter of the inference compute; the weights are due later this month. A Codex update makes GPT-6 Astra and GPT-6.1 Sol roughly 50% faster by default, while a UT Austin study covering about 35,000 runs found that reducing context to roughly one-third of the tokens could still make runs 20%–80% slower. OpenAI will add statistical watermarks to eligible ChatGPT and Codex text in the EU, but rewriting or translation can remove the watermark, and the detector will initially be available only to approved researchers.
My take
My takeaway today is that efficiency gains can look very different when you change the yardstick. Context compression can actually slow things down, and Claude’s subscription advantage shrinks when measured by task cost; I care more about the time and cost of completing comparable tasks than token savings or generous allowances.
Beam targets inference compute efficiency, while Rho-1 brings robotic actions into an omnimodal model.
Reflection has unveiled its first open-weight model, Beam, which uses a sparse MoE architecture with 501 billion total parameters and 23 billion active parameters, targeting coding, reasoning, and agent tasks. The company says Beam's advanced reasoning scores approach those of GLM-5.2 while using just one-third to one-quarter of its inference compute. The model is still undergoing final red-team testing and evaluation. Applications for early access are now open, with the weights and technical report planned for release later this month.
Read original →Reka has released a research preview of Rho-1, an omnimodal model with 19 billion parameters. It understands and generates text, images, video, and robotic actions within a single neural network, bringing multimedia capabilities and robotic action generation into one model. Reka says the model was trained from scratch in roughly three months using 320 H100 GPUs.
Read original →🔍 Analysis: Beam's comparison with GLM-5.2 uses inference compute as its yardstick, while Rho-1 puts robotic actions and multimedia generation into a single network. The two cannot be ranked solely by parameter count. For developers, model selection starts with distinguishing the goals: for coding and agent tasks, consider task performance per unit of compute; for omnimodal applications, consider whether cross-modal understanding translates into usable generated outputs.
Model speedups, subscription allowances, and context compression need to be compared on the same task basis.
Tibo announced a day-one update: GPT-6 Astra and GPT-6.1 Sol are now roughly 50% faster by default. The speedup applies to subscription usage across all of its products, as well as partner products accessed through Sign in with ChatGPT, including OpenCode, Pi, Amp, and Devin. Users do not need to change any settings to benefit.
Read original →SemiAnalysis tested subscription plans from Anthropic, OpenAI, and other providers, finding that Claude offered more than five times OpenAI's API-equivalent value. The tests estimated allowance consumption by isolating different token types and observing changes in usage meters; the value of the same plan varies with the model and workload. After adjusting for task cost, scaling01 narrowed Claude's advantage to 1.3–2.9 times.
Read original →UT Austin studied context compression for coding agents across roughly 35,000 runs. The results show that some compression methods use only about one-third of the tokens in the full context yet can run 20%–80% slower; reducing tokens does not necessarily reduce latency. Threshold-triggered compression outperformed step-triggered compression, but the best strategy varied by model.
Read original →🔍 Analysis: The roughly 50% speedup for GPT-6 Astra and GPT-6.1 Sol cannot simply be added to the token savings from context compression: the 20%–80% slowdown measured by UT Austin shows that compression strategies can offset speed gains. Claude's subscription advantage also shrinks from more than five times the API-equivalent value to 1.3–2.9 times when measured by task cost, underscoring that token allowances are not the same as final output. For businesses, a more useful basis for comparison is the time and cost required to complete comparable tasks, rather than speed or allowances in isolation.
Cognition manages long-term memory, while Cursor adds a way to steer agents during execution.
Cognition has released the git- and Markdown-based Agent Memory Repo as an open standard, accompanied by Dreaming, a memory-maintenance agent that runs periodically. It identifies patterns in cross-session records and adds memories, while merging duplicate entries, deleting outdated content, and checking sources to resolve contradictions. Devin's local trial requires manual invocation of memory features and does not include automatic startup or scheduled Dreaming.
Read original →The Cursor SDK now supports mid-run steering, allowing developers to send messages while an agent is executing. After a call to run.steer(), the new message is added to the agent's next interaction. If a subagent is executing a task at that point, it moves to the background and continues working.
Read original →🔍 Analysis: Cognition's Agent Memory Repo determines which past experiences carry into later sessions, while Cursor's run.steer() feeds new instructions into the next interaction. They address long-term state and course correction during current execution, respectively. Adding these capabilities does not create an automatic closed loop: memory in Devin's local trial still requires manual invocation, and subagents can keep working in the background while Cursor is being steered. Developers therefore need to account for both the triggers for memory maintenance and the state of tasks still running in their control logic.
EU product coverage, the optional global API toggle, and detector access each have distinct limits.
OpenAI will add invisible statistical watermarks to eligible text generated by ChatGPT and Codex in the EU, while offering an optional toggle for the API globally. The move responds to EU regulatory requirements, extending existing image and audio provenance measures to text. However, rewriting or translation can remove the watermark, and the detector will initially be available only to approved researchers.
Read original →More than 90 Opus 5.5 agents identified two simulated candidates in three days.
Vals AI says more than 90 Opus 5.5 agents helped identify two candidate room-temperature magnetic semiconductors in three days: YBaMnFeO₅ and KV[Cr(CN)₆]. The exploration addressed a decades-long goal in materials research: sorting electrons by spin while having the material's magnetic contributions cancel each other out. Both materials remain candidates identified in simulations.
Read original →