Fireworks says ARCv3 reduces BF16 weight-update delta payloads to an average 0.19% of original weight size, compared with 0.36% for ARCv2, while reconstructing the trainer’s exact weights. It is the default for Fireworks Trainer SDK rollouts; teams bringing their own trainer can integrate it through the fireworks-delta-compression package.
Fireworks reports that an oracle choosing among 18 models per DeepSWE task reached a 97.6% pass rate at $1.88 per task, versus 74.1% at $6.52 for GPT-6 Astra. The result is hindsight-based and evaluated on the same 113 tasks used to choose winners, so it is not evidence of a production router achieving that performance.
Fireworks says FireRouter with Opus is available as a standalone router model and through its CLI, including in Claude Code, Codex and Cursor. It reports lower cost and slightly lower accuracy than Opus-only in an internal coding test.
OpenRouter outlines separate checks for tool selection, argument structure and values, and multi-step call trajectories. It recommends keeping cases, graders, settings, and routing consistent when comparing models.
OpenRouter’s asynchronous Batch API lets providers complete requests within a 24-hour window in exchange for generally charging 50% or less of normal per-token prices. It supports chat completions, responses, messages, and embeddings, with timing varying by batch and submission hour.
OpenRouter reports Jev 1.13 scored 81.0% accuracy versus Claude Opus 5 at 84.4% on 3,080 Banking77 utterances. With prompt caching, Jev cost $0.11 per 1,000 requests versus $2.42 for Opus; the post also tests confidence-based routing and notes important limits to the comparison.
Google says Gemini 3.8 Live with Live Avatar is available in Gemini Enterprise. Organizations can also request custom avatar creation through enterprise allowlisting, while API documentation is provided for getting started.
GitHub says agentic autofix can use existing Copilot Memories for security-alert context and save fix patterns as memories for future use. Both features are in public preview.
xAI says Team Bots let Teams and Enterprise users share Grok Bots with common context, plugins, credentials, and team skills while keeping individual conversations and memories separate. Bots can also be used collaboratively in Slack.
xAI says Grok Voice Transcribe 2.0 improves accuracy over version 1.0 without changing its Speech-to-Text API prices. Batch transcription costs $0.10 per audio hour and streaming costs $0.20 per audio hour; xAI says the new model will soon become the default and version 1.0 will be deprecated in the coming weeks.
Grok Build now records project conventions, decisions, and facts in the background, then reads relevant notes in later sessions. Notes are stored as markdown and can be browsed or organized with commands.
GitHub says GPT-6.1 Sol is generally available to Copilot Pro+, Max, Business, and Enterprise users, with gradual rollout across listed Copilot surfaces. Usage-based billing follows provider list pricing; Business and Enterprise administrators can manage access in Copilot settings.
xAI says Grok 4.7 is available in Cursor and Grok Build, as well as through its API and other listed channels. API pricing starts at $2 per million input tokens and $6 per million output tokens; a faster variant costs twice as much.
Cursor says Rollouts monitors changes from pull request through production and can alert, pause a progressive rollout, or prepare a revert PR for approval. Security Reviewer checks pull requests for vulnerabilities and proposes fixes. Both are available on Teams and Enterprise plans.
OpenAI lists GPT-6.1 Sol for Responses and Chat Completions, with standard pricing for prompts up to 272K input tokens of $2 per million input tokens, $0.10 per million cached input tokens, $2.50 per million cache-write tokens, and $10 per million output tokens. Multi-agent support is in beta through a Responses API request.
Anthropic says Claude Sonnet 5.5 is available on several named API and cloud platforms. It also documents compatibility changes affecting thinking settings, forced tool use, thinking blocks, computer use, and advisor-tool requests.
Anthropic lists Opus 5.5 API input and output at $4 and $20 per million tokens, below Opus 5 prices. It also says five-hour usage limits are increasing on Pro, Max, Team, and seat-based Enterprise plans, with a reset subscribers can save and use when they choose.
GitHub says Claude Sonnet 5.5 is generally available to Copilot Pro, Pro+, Max, Business, and Enterprise users, with gradual rollout across listed Copilot surfaces. Business and Enterprise administrators can manage model access in settings.
OpenAI says GPT-6 Sol and GPT-6 Luna accept text and image inputs and generate text through the Responses and Chat Completions APIs. Standard pricing for prompts up to 272K input tokens is $2 input, $0.20 cached input, and $10 output per million tokens for Sol; Luna is $0.10, $0.01, and $0.50 respectively.
Anthropic says cache diagnostics are out of beta on the Claude API. Requests can opt in with a diagnostics object, and Messages responses now include a diagnostics field that is null when diagnostics were not requested.