SessionWatcher guides

AI news and practical updates

What changed, what the official sources say, and what it means for your workflow.

Published by SessionWatcher. AI assisted reporting with a separate factual review. Read our editorial process.

News ·

Fireworks releases ARCv3 for smaller lossless RL weight updates

Fireworks says ARCv3 reduces BF16 weight-update delta payloads to an average 0.19% of original weight size, compared with 0.36% for ARCv2, while reconstructing the trainer’s exact weights. It is the default for Fireworks Trainer SDK rollouts; teams bringing their own trainer can integrate it through the fireworks-delta-compression package.

News ·

Fireworks reports model routing scores 97.6% on DeepSWE at $1.88 per task

Fireworks reports that an oracle choosing among 18 models per DeepSWE task reached a 97.6% pass rate at $1.88 per task, versus 74.1% at $6.52 for GPT-6 Astra. The result is hindsight-based and evaluated on the same 113 tasks used to choose winners, so it is not evidence of a production router achieving that performance.

News ·

Fireworks launches FireRouter with Opus for coding harnesses

Fireworks says FireRouter with Opus is available as a standalone router model and through its CLI, including in Claude Code, Codex and Cursor. It reports lower cost and slightly lower accuracy than Opus-only in an internal coding test.

News ·

How to test tool-calling accuracy in AI agents

OpenRouter outlines separate checks for tool selection, argument structure and values, and multi-step call trajectories. It recommends keeping cases, graders, settings, and routing consistent when comparing models.

News ·

OpenRouter adds Batch API with generally 50% lower per-token pricing

OpenRouter’s asynchronous Batch API lets providers complete requests within a 24-hour window in exchange for generally charging 50% or less of normal per-token prices. It supports chat completions, responses, messages, and embeddings, with timing varying by batch and submission hour.

News ·

OpenRouter compares Jev 1.13 and Claude Opus 5 on Banking77 classification

OpenRouter reports Jev 1.13 scored 81.0% accuracy versus Claude Opus 5 at 84.4% on 3,080 Banking77 utterances. With prompt caching, Jev cost $0.11 per 1,000 requests versus $2.42 for Opus; the post also tests confidence-based routing and notes important limits to the comparison.

News ·

xAI releases Grok Voice Transcribe 2.0 at unchanged API prices

xAI says Grok Voice Transcribe 2.0 improves accuracy over version 1.0 without changing its Speech-to-Text API prices. Batch transcription costs $0.10 per audio hour and streaming costs $0.20 per audio hour; xAI says the new model will soon become the default and version 1.0 will be deprecated in the coming weeks.

News ·

Grok Build adds memory across coding sessions

Grok Build now records project conventions, decisions, and facts in the background, then reads relevant notes in later sessions. Notes are stored as markdown and can be browsed or organized with commands.

News ·

GPT-6.1 Sol becomes generally available in GitHub Copilot

GitHub says GPT-6.1 Sol is generally available to Copilot Pro+, Max, Business, and Enterprise users, with gradual rollout across listed Copilot surfaces. Usage-based billing follows provider list pricing; Business and Enterprise administrators can manage access in Copilot settings.

News ·

Grok 4.7 launches in Cursor, Grok Build and through the API

xAI says Grok 4.7 is available in Cursor and Grok Build, as well as through its API and other listed channels. API pricing starts at $2 per million input tokens and $6 per million output tokens; a faster variant costs twice as much.

News ·

Cursor adds Rollouts and Security Reviewer bots to Teams and Enterprise

Cursor says Rollouts monitors changes from pull request through production and can alert, pause a progressive rollout, or prepare a revert PR for approval. Security Reviewer checks pull requests for vulnerabilities and proposes fixes. Both are available on Teams and Enterprise plans.

News ·

OpenAI releases GPT-6.1 Sol with beta Responses API multi-agent support

OpenAI lists GPT-6.1 Sol for Responses and Chat Completions, with standard pricing for prompts up to 272K input tokens of $2 per million input tokens, $0.10 per million cached input tokens, $2.50 per million cache-write tokens, and $10 per million output tokens. Multi-agent support is in beta through a Responses API request.

News ·

Claude Sonnet 5.5 becomes generally available in GitHub Copilot

GitHub says Claude Sonnet 5.5 is generally available to Copilot Pro, Pro+, Max, Business, and Enterprise users, with gradual rollout across listed Copilot surfaces. Business and Enterprise administrators can manage model access in settings.

News ·

OpenAI releases GPT-6 Sol and GPT-6 Luna for Responses and Chat Completions

OpenAI says GPT-6 Sol and GPT-6 Luna accept text and image inputs and generate text through the Responses and Chat Completions APIs. Standard pricing for prompts up to 272K input tokens is $2 input, $0.20 cached input, and $10 output per million tokens for Sol; Luna is $0.10, $0.01, and $0.50 respectively.