ai-powered-markdown-translatorArticle translated from fr to en with gpt-6.1-sol.
Mistral releases Mistral Large 4 in preview, a multimodal model with 1 trillion parameters whose API is available to everyone starting today and whose weights are promised for late October. Anthropic brings Claude into Google Docs, Sheets and Slides and restructures its cyber verification program into three access tiers, while Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model that fits on a phone. Decision models also have their moment: OpenAI opens its Decisions API in public beta, Perplexity halves its price, and a community ranking compares them.
Mistral Large 4: 1 trillion parameters in preview, weights promised for late October
October 6 — Mistral AI launches Mistral Large 4 in public preview, nicknamed ML4 and, “very officially,” the Chonk. It is the company’s largest model to date: a natively multimodal mixture of experts (Mixture-of-Experts) with 1 trillion parameters (1T), including 49 billion active parameters according to the announcement post, combining instruction following, reasoning and agentic capabilities within a context of one million tokens. The documentation page gives different figures without explaining the discrepancy: 1.05 trillion parameters, including 52 billion active parameters, plus a 1.6 billion vision encoder.
The API is available to everyone starting today on Mistral Studio, but the weights have not yet been released: Mistral promises them by the end of October, along with details on the architecture and post-training. In the meantime, cybersecurity leaders, verified partners and government authorities are testing the model under real-world conditions with reduced moderation. ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s European data centers, which also serve the preview. The company plans a European deployment that it operates from end to end, under European law, and presents the model as the first milestone in the roadmap funded by its €3 billion Series D.
Cybersecurity is the main selling point. According to Mistral, ML4 ranks among the top five models in the Artificial Analysis Cyber Index and scores 82 % on one of its tests, which involves reproducing a real vulnerability in open source software and then fixing it, the highest score of any model. The company says Claude Opus 5.5 and GPT-6 Astra score close to zero on the same test because they refuse the task.
| Benchmark evaluated | Score reported by Mistral | Comparison cited by Mistral |
|---|---|---|
| DeepSWE v1.1 | 61,7 % | — |
| Coding Agent Index (combined) | 49,8 % | ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max |
| Surge AI blind human evaluation (code, out of 5) | 3,74, 2nd out of 5 | behind Claude Opus 5 (4,22) |
| Cybench (40 challenges) | 93 % | — |
| AA Cyber Index, reproducing then fixing a vulnerability | 82 % | highest score of any model |
| AutomationBench (657 business processes) | 59,9 % | ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro |
The post and the documentation page also differ on pricing: the card in the post lists $1.36 per million input tokens and $4.18 for output, while the documentation page strikes through those rates and displays a price cut in half alongside them, with no duration or explanation.
| ML4 characteristic | According to the announcement post | According to the documentation page |
|---|---|---|
| Total parameters | 1 000 billion | 1 050 billion |
| Active parameters | 49 billion | 52 billion |
| Input, per million tokens | 1,36 dollar | 0,68 dollar (1,36 struck through) |
| Output, per million tokens | 4,18 dollars | 2,09 dollars (4,18 struck through) |
All these scores come from Mistral, which says reinforcement learning for the preview is continuing with no sign of saturation and expects rapid progress in the coming weeks.
🔗 Mistral’s post on Mistral Large 4 🔗 Mistral Large 4 documentation page
Claude for Google Workspace: Claude comes to Docs, Sheets and Slides
October 6 — Anthropic launches Claude for Google Workspace, an add-on (add-on) that opens Claude in a side panel in Google Docs, Sheets and Slides, in public beta on all paid plans. Claude reads the open file, sees the current selection (text, cells or slides) and edits the file in place.
| Google application | What Claude does in it |
|---|---|
| Docs | In-place corrections and style changes, suggestion cards to apply or dismiss for more substantial rewrites |
| Sheets | Formulas, pivot tables, native charts, new tabs, processing a range through Python for a join or cleanup |
| Slides | Slides based on the presentation’s layouts and theme, checking for overlapping or overflowing elements |
Users choose the level of autonomy: in “ask before edits” mode (Ask before edits), enabled by default, every change goes through an approval card; in “accept all edits” mode (Accept all edits), Claude keeps going without stopping. Connected to the Claude account, the panel uses the same models, connectors and skills (skills). On Enterprise plans, the Compliance API, customer-managed encryption keys (CMEK) and OpenTelemetry audit export also apply to the add-on.
Claude now works inside Google Docs, Sheets, and Slides, and those files also open inside Claude. In Google Workspace, Claude sits in a sidebar next to your file, reads what you have open, and edits it in place. You can approve each edit before it lands. — @claudeai on X
The reverse direction arrives at the same time: new Google Docs, Sheets and Slides connectors, also in beta, let users create and edit Google files from Claude by pasting a link or requesting a new document. On supported configurations, the file opens in a panel alongside the conversation, and Claude’s access follows Google’s sharing permissions. The add-on installs from the Google Workspace Marketplace (Extensions > Claude > Open Claude), and administrators can deploy it from the Google Admin console; on Team and Enterprise, an owner must first enable the connectors. Google Workspace thus joins Microsoft 365, where Claude for Excel, PowerPoint and Word has been generally available since May 7.
🔗 Anthropic’s post on Claude for Google Workspace
Cyber Verification Program: Anthropic opens access to Opus 5.5, Sonnet 5.5 and Mythos 5.1 through three access tiers
October 6 — Anthropic restructures its cyber verification program (Cyber Verification Program, CVP), which gives verified security professionals access to capabilities that its consumer models block. Two initiatives had coexisted for six months: Project Glasswing, which gave organizations responsible for the most critical software access to Claude Mythos, and a single-tier CVP that relaxed the guardrails on Opus and Sonnet models. They merge into three tiers, all of which provide access to Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1 and upcoming models. Generally available models (Opus 5.5, Fable 5.1, Sonnet 5.5) retain conservative guardrails that block most cybersecurity work.
| Access tier | Uses covered | Intended audience and timeframe |
|---|---|---|
| Defense Access | Security operations center, incident response, malware reverse engineering, vulnerability analysis | Security teams, critical infrastructure of any size, open source maintainers, researchers; a few days |
| Red Team Access | Defensive uses, plus authorized penetration testing and simulated attack exercises (red teaming) | Organizations only, on systems they are authorized to test; a few weeks |
| Specialized Access | Testing safety systems: aviation, power grids, telecoms, interbank transfers, government networks | A few organizations reviewed individually with the US government; Glasswing members transferred |
In Red Team Access, actions that could cause physical harm or massive disruption, such as deploying ransomware, remain blocked in real time. To verify its settings, Anthropic ran Opus 5.5 through CyScenarioBench, an evaluation of multistep cyber operations, with the guardrails of each tier (10 challenges, 5 attempts each).
| Configuration tested by Anthropic (Opus 5.5) | CyScenarioBench result (50 attempts) |
|---|---|
| Without CVP | All tasks blocked at the first prompt |
| Defense Access | 46 attempts blocked at some point, 4 successful |
| Red Team Access | No blocks, 34 tasks completed, equivalent to the unguarded model’s 67,6 % |
The post also offers an initial assessment of Glasswing, with partial figures according to Anthropic: at least 129,000 verified vulnerabilities found by partners from April to July, plus 5,500 from Anthropic’s open source analyses, including more than 33,000 classified as critical or high. These figures are based on 33 partner reports, and Anthropic estimates the actual impact to be “at least five times higher.” In a testimonial published the same day, Comcast says it found a critical authentication vulnerability with Claude Mythos Preview while evaluating 258 systems and roughly 170 million lines of code, and a Booz Allen analyst reviewed 138 repositories in twelve days.
On the practical side, the program requires data retention to monitor abuse, pending Enterprise Frontier Safeguards (EFS), which will allow this data to remain in a customer-controlled cloud later this fall; organizations already using Fable 5.1 or Mythos 5.1 with zero data retention can do the same with the CVP. The program is available on Claude Platform, Vertex AI and Microsoft Foundry, and on Amazon Bedrock only for customers eligible for EFS. Individuals can apply only for Defense Access, on a paid plan, and this tier will have to adopt phishing-resistant multifactor authentication and stop using API keys by December 15. A webinar is scheduled for October 14 at 9 a.m. PT.
🔗 Anthropic’s post on the Cyber Verification Program 🔗 Cyber Verification Program help page 🔗 Comcast and Booz Allen with Claude Mythos
EmbeddingGemma 2: Google’s open embedding model becomes multimodal
October 6 — Google DeepMind releases EmbeddingGemma 2, the second generation of its open embedding model designed to run on-device. The first EmbeddingGemma, downloaded more than 20 million times according to Google, handled only text; the new model maps text (including code), images, video and audio, individually or in combination, into the same 768-dimensional space. Built on the Gemma 4 architecture, it is released under the Apache 2.0 license.
Meet EmbeddingGemma 2, our first natively multimodal open model for on-device embeddings. It expands beyond text to unify code, images, audio, and video in a shared space. — @GoogleDeepMind on X
The model has 740 million parameters, but it is modular: a 270M text core, with a vision encoder (170M) and an audio encoder (300M) added as needed. A text-only application therefore loads only 270M parameters. With quantization, Google measures around 191 Mo of active RAM for the text weights and 567 Mo for the complete multimodal model on a Pixel 11 Pro.
| EmbeddingGemma 2 feature | Value reported by Google |
|---|---|
| Total parameters | 740M (text 270M, vision 170M, audio 300M) |
| Vector dimensions | 768, truncatable to 512, 256 or 128 |
| Context window | 8K tokens, four times that of the first version |
| Active RAM on Pixel 11 Pro (quantized) | around 191 Mo for text only, 567 Mo for multimodal |
| MTEB Code | 78,68, compared with 68,76 for the first version |
| Multilingual MTEB v2 | 61,36, compared with 61,15 for the first version |
“Russian doll” representation learning (Matryoshka Representation Learning) allows vectors to be shortened, reducing storage requirements by up to a factor of six; according to the model card, quality is largely unaffected down to 256 dimensions, while 128 dimensions are mainly suitable for text-only use cases. The clearest improvement is in code, with a gain of nearly ten points on MTEB Code according to Google, while multilingual text remains at the previous level.
The intended uses are entirely local search and retrieval-augmented generation (RAG), such as finding a passage in a video from a voice memo. Since the model shares Gemma 4’s text tokenizer (tokenizer) and audio encoder, the two can run together with a smaller memory footprint. The weights are available on Hugging Face and Kaggle; availability in the Model Garden of Gemini Enterprise Agent Platform is announced as “coming soon,” with no date. The model is supported at launch by Transformers (version 5.19.0), sentence-transformers, MLX, vLLM, llama.cpp, Ollama and LM Studio, as well as in the browser with transformers.js.
🔗 Announcement on the Google blog 🔗 EmbeddingGemma 2 model card 🔗 Weights on Hugging Face
Decision models: OpenAI opens its Decisions API, Perplexity releases Decider v1.1 and halves its price, Decision Index 0.3 ranks them
Decision models, which choose an option or score an input instead of writing a free-form response, had a busy day: two providers revise their offerings and pricing, and the community leaderboard for the category changes its methodology.
OpenAI’s Decisions API enters public beta
October 6 — A week after opening limited access at DevDay, OpenAI’s Decisions API enters public beta, with general availability expected “in the coming weeks.” It answers closed-ended questions about text, an image or both, using GPT-6 Luna, the only available model, through a dedicated endpoint (POST /v1/decisions); according to OpenAI, its typed responses arrive around 10 times faster than with the Responses API. Today’s new detail is pricing: 0,10 dollars per million input tokens, with no charges for output or caching, plus additional fees for regional processing and long contexts.
| Question type | Intended use | Returned result |
|---|---|---|
Predicate (predicate) | Check a condition, such as a damaged product in a photo | Probability from 0 to 1 that the condition is true |
Choice (choice) | Select an option from a supplied list | The option, each option’s probability and a confidence score |
Score (score) | Rate on ordered levels, such as the severity of an incident | Average of the levels weighted by their probabilities |
Several questions can address the same input in a single request, and images must be sent in base64, without a hosted URL or file ID. The API supports Zero Data Retention (Zero Data Retention) and HIPAA use cases for eligible customers, with data residency in the United States and Europe.
🔗 Decisions API guide 🔗 OpenAI API changelog
Perplexity’s pplx-decider-v1.1-27b and the price cut in half
October 6 — Five days after launching its own Decisions API, Perplexity updates the model powering it, in a thread from its developer account, @perplexitydevs. Released with open weights under the Apache 2.0 license on Hugging Face, pplx-decider-v1.1-27b retains v1’s Qwen3.8-27B base, reads text and images within a 250k-token context and removes the causal mask (causal mask) from its full-attention layers. The Decisions API, which now serves this model, costs 0,02 dollars per million input tokens, half v1’s 0,04 dollars, and output remains free.
| Decision Index category (Hugging Face model card) | Jev score | v1 score | v1.1 score |
|---|---|---|---|
| Knowledge (Knowledge) | 51,4 | 40,9 | 48,18 |
| Language (Language) | 62,0 | 63,5 | 69,45 |
| Information retrieval (Retrieval) | 55,4 | 54,9 | 61,26 |
| Tools (Tools) | 75,1 | 79,3 | 78,88 |
| Arts (Arts) | 37,7 | 39,4 | 44,66 |
| Weighted overall score | 57,9 | 56,4 | 61,56 |
According to the model card, the overall score rises from 56,4 to 61,56, ahead of the 57,9 scored by Jev, TypeSafe AI’s decision model, which nevertheless retains its lead in knowledge. Perplexity also claims the highest score on the new Decision Index 0.3. Self-hosting requires a GPU capable of holding around 49 Gio of weights and the supplied implementation, because standard causal inference does not reproduce the measured behavior.
🔗 @perplexitydevs’ thread on X 🔗 pplx-decider-v1.1-27b on Hugging Face
Decision Index 0.3
October 6 — apolinario, an engineer at Hugging Face, releases version 0.3 of the Decision Index, the community leaderboard of open decision models that reproduce Jev’s Decision Model. It now includes 112 models, including 111 open reproductions, running on an NVIDIA RTX PRO 6000. The methodology changes: the full score (Full score) combines 37 public benchmarks, private tests measuring the same skills and private tasks from new domains, so that rankings reflect skills rather than just performance on known benchmarks. A vision tab also makes its debut.
| Displayed rank | Ranked model | Full score |
|---|---|---|
| 1 | Perplexity Decider v1.1 (27B) | 62,8 |
| 2 | Fastino GLiDE without reasoning (28B) | 60,2 |
| Reference | Jev | 60,1 |
| 3 | Torchcast Decision 27B | 59,9 |
| 4 | deck31b (Gemma 4 31B) | 59,0 |
Decider v1.1’s two scores should not be confused: 61,56 is the score published in Perplexity’s model card, while 62,8 is the score calculated by this leaderboard using its new methodology. The private tests, by their nature, are not published.
🔗 apolinario’s announcement on X 🔗 Decision Index 0.3
Computer use: OpenAI trains GPT-6 Astra on Ironclad contracts
October 6 — OpenAI launches research collaborations with software vendors to make its agents more effective in business software, a computer use initiative (computer use), starting with Ironclad, a specialist in AI-assisted contracts. The teams defined 11 legal, sales and procurement tasks, each taking an experienced user 30 to 40 minutes according to OpenAI, evaluated against 8 to 50 criteria. Ironclad provided hosted environments of its software, where the models trained through reinforcement learning on synthetic tasks drawn from public contracts in the SEC’s EDGAR database.
| OpenAI’s measurement across the 11 tasks | GPT-5.6 Sol (High reasoning) | GPT-6 Astra (Max reasoning) |
|---|---|---|
| Average score | 41,6 % | 55,0 % (+32 %) |
| Estimated average time per attempt (simulated) | 37,0 minutes | 19,2 minutes (-48 %) |
GPT-6 Astra is OpenAI’s first frontier model trained on these tasks, and an internal model used during its development reaches 63,7 %. These figures come from OpenAI, and the times are simulated using assumed processing speeds, rather than measured at customer sites. OpenAI invites other vendors to propose tasks that current agents still cannot reliably complete.
🔗 OpenAI’s post on its collaboration with Ironclad
APIs and platforms: OpenAI API tiers and BAA, documents and images in Perplexity’s Agent API, GLM-5.3 on Bedrock
Four changes affect developers buying inference: access requirements for OpenAI’s API, tools in Perplexity’s Agent API and a new open model in AWS’s catalog.
OpenAI API usage tiers go from five to three
October 6 — OpenAI simplifies access to its API’s highest rate limits (rate limits): the five paid usage tiers become three, Build, Launch and Grow. The highest, Grow, becomes available after 500 dollars in cumulative API payments, compared with 1 000 dollars for the previous highest tier. Organizations already on a paid tier switch automatically, with no action required, then move up tiers when their cumulative credit purchases reach the thresholds; the free tier remains capped at 100 dollars per month.
| Usage tier | Cumulative purchase threshold | Monthly usage cap | Astra, Sol and Terra (requests and tokens per minute) | Luna (requests and tokens per minute) |
|---|---|---|---|---|
| Build | 5 dollars | 500 dollars | 5 000 and 1 000 000 | 5 000 and 2 000 000 |
| Launch | 100 dollars | 5 000 dollars | 10 000 and 4 000 000 | 10 000 and 10 000 000 |
| Grow | 500 dollars | 200 000 dollars | 15 000 and 40 000 000 | 30 000 and 180 000 000 |
🔗 @OpenAIDevs’ thread on X 🔗 Rate limit documentation
Self-service BAA and HIPAA compliance on the OpenAI API
October 5 — Organizations processing protected health information with OpenAI’s API can now sign the business associate agreement (Business Associate Agreement, BAA) required by the US HIPAA law themselves. In Settings > Organization > General, an administrator accepts the standard BAA and enables HIPAA compliance support, without an enterprise contract. This self-service option is limited to eligible organizations with an established history of API usage, and activation cannot be undone from these settings. Custom terms still require an email request, with a response within one to two business days. OpenAI reminds users that signing the BAA alone does not make an application compliant, and that no BAA is offered for ChatGPT Business.
🔗 OpenAI help article on the API BAA
Perplexity’s Agent API reads documents and searches for images
October 5 — The Perplexity API changelog receives three entries late in the evening. The Agent API now accepts PDF, DOC, DOCX, TXT and RTF documents as input, embedded in the request or referenced by a public HTTPS URL, with support depending on the selected model and provider. A new tool, image_search, searches the web for images and returns structured results, including the image URL and the URL of its source page: from 1 to 30 images per call, 5 by default, domain and format filters, and safe search enabled by default. It costs 2,50 dollars per 1 000 successful invocations, and failed calls are not charged. The third entry corrects cost accounting: responses now detail the cost of extractions performed in the sandbox, at GPT-6 Luna’s unchanged rate.
🔗 Perplexity API changelog 🔗 image_search tool documentation
Z.ai’s GLM-5.3 on Amazon Bedrock
October 6 — Z.ai announces the arrival of GLM-5.3, its flagship open-weight model, on Amazon Bedrock; the AWS model card dates the launch to October 5. It is a mixture of experts with 744 billion parameters, around 40 billion of them active per token, with a context of one million tokens and up to 128 000 output tokens. Access is limited to eligible customers, who go through their AWS account team, and the model is served only through cross-region inference (US or global profile), with prompt caching starting at 1 024 tokens and Standard, Priority, Flex and Reserved service tiers. The model card gives no price and refers readers to the Bedrock pricing page. GLM-5.3 follows Kimi K3 (September 22) and Grok 4.7 (September 28), which arrived on Bedrock over the past two weeks.
🔗 GLM 5.3 model card on Amazon Bedrock 🔗 Z.ai’s announcement on X
Coding agents: Claude Code 2.1.290 to 2.1.292, cloud sessions, Vibe CLI 2.26.0, Antigravity and Replit
Coding agents receive five updates, from the terminal to the editor: three Claude Code releases in less than twenty hours, a guide to cloud sessions, a Vibe CLI whose new Rust interface gains more features, the Antigravity extension for VS Code and Replit, which can see all of the user’s projects.
Claude Code 2.1.290 to 2.1.292
October 5 and 6 — Released on October 5 at 23:33 UTC, Claude Code 2.1.290 is a substantial release: 190 entries, including 131 fixes. claude attach and claude logs accept part of a session’s name instead of its identifier, and /claude-api managed-agents-onboard turns the Managed Agents template described on a web page into ant apply files. The interactive session’s web search budget no longer stops after 200 calls: it replenishes at a rate of 100 calls per hour. At the medium effort level (medium), /code-review also flags deviations from CLAUDE.md conventions on Opus 5.5 and Sonnet 5.5. In Slack, Claude Tag gains a fast mode (!fast), a Claude in Chrome browser_batch call gets 90 seconds instead of 60, and WebFetch no longer silently truncates text beyond 100,000 characters. pyright and more forms of ps require permission, and several permission loopholes are fixed: shell-expanded rg and git grep wildcards, CLAUDE.md files linked outside working directories, and an organization mod bypassed by a user mod. Four hours later, 2.1.291 fixes two regressions, one introduced in 2.1.290 (responses to permission prompts lost in cloud sessions), the other in 2.1.288 (the last messages of a session lost on exit).
On October 6 at 18:59 UTC, 2.1.292 (92 entries, including 64 fixes) adds an effort parameter to the Agent tool to launch a subagent at the requested effort level, and claude plugin install --marketplace, which adds the marketplace (marketplace) if needed and then installs the plugin. Local MCP servers (stdio) now negotiate protocol 2026-07-28 by default (MCP_PROTOCOL_NEGOTIATION=legacy to revert), CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS extends retries for 529 errors, and mods gain a prompt autocompletion event and prompt caching. On the security side, approvals from a PreToolUse hook and auto mode no longer bypass the permission prompt for reads from network paths (UNC), and <system-reminder> tags written by a hook are escaped before reaching Claude. The Artifact tool finally lists 200 artifacts instead of 50, and Code Review breaks down its statistics by repository. Neither 2.1.290 nor 2.1.292 was announced on X.
| Claude Code version | Release date (UTC) | Number of entries | Main new feature |
|---|---|---|---|
| 2.1.290 | October 5, 23 h 33 | 190, including 131 fixes | Sessions identified by name, managed-agents-onboard, replenished WebSearch budget |
| 2.1.291 | October 6, 3 h 55 | 2 fixes | Two regressions fixed |
| 2.1.292 | October 6, 18 h 59 | 92, including 64 fixes | Subagent effort, plugin and marketplace in one command, MCP 2026-07-28 |
🔗 Claude Code 2.1.290 🔗 Claude Code 2.1.291 🔗 Claude Code 2.1.292
The guide to Claude Code cloud sessions
October 6 — Two weeks after Claude Code cloud sessions entered preview, Anthropic publishes a field guide by Addy Osmani on claude.dev. Each task runs on a fresh virtual machine with roughly 4 vCPU, 16 GB of RAM and 30 GB of disk space, with the repository cloned onto a new branch, at no additional compute cost on Pro, Max, Team and Enterprise plans. The guide describes seven use cases: clearing a list of pending tasks (backlog) in parallel, having a fix proven through repeated runs, planning locally then resuming the session with claude --teleport, monitoring from a phone, handing CI failures and review comments to Auto-fix, triggering routines and running potentially unsafe code in a disposable VM. In his demonstration, all three sessions finished 87 seconds after the first launch, and a project can launch up to 200 new conversations (threads) per day. The guide also details the GitHub connection: signing in with GitHub and installing the Claude GitHub app are two separate permissions, and only the latter grants access to private repositories. The $100 (Pro) or $250 (Max) bonus credit must be claimed before October 7 at 23:59 PT and expires on November 4.
🔗 The cloud sessions field guide on claude.dev 🔗 @ClaudeDevs’ reminder about the bonus credit
Vibe CLI 2.26.0
October 6 — Mistral releases Vibe CLI 2.26.0, its command-line coding agent, thirteen days after 2.25.8 and without an announcement on X. With 33 additions, 23 changes and 91 fixes, the update is substantial: 72 of its 147 entries concern the new interface written in Rust: a first-launch wizard, voice mode (/voice), recurring prompts described in natural language with /loop, sending the session to Vibe Code Web with /teleport and the vibe update command. Vibe now always runs on the Unified Harness, its new runtime engine, and stops with an explicit error if it is missing, instead of silently falling back to the old engine. Plugins can bundle their own MCP server, the fallback API key stored in ~/.vibe/.env is now readable only by its owner, and the Rust CLI now sends its usage telemetry (startup time, commands used, terminal detected) to Mistral.
🔗 Mistral Vibe v2.26.0 release notes
The Antigravity extension for VS Code 1.7.0
October 5 — Google releases version 1.7.0 of the Antigravity extension for Visual Studio Code, which brings Google Antigravity agents into Microsoft’s editor (side-panel conversation, inline diff review, interactive plans), starting with VS Code 1.90. Updated roughly weekly since early September according to its changelog, it had never been covered here. This 1.7.0 release (6 improvements, 9 fixes) improves code review: Accept and Reject decisions can be undone and redone with the keyboard, the side-by-side diff editor gains Accept All and Reject All buttons, and VS Code’s automatic saving (Auto Save) can no longer overwrite or close an ongoing review. On Windows, native notifications signal task completion when the window is not in the foreground, and Google Cloud Workstations and Cloud Shell gain interactive authentication. On the same day, the Visual Studio extension moves to 1.0.261005.0 and attaches diagnostic logs to submitted feedback.
🔗 Google Antigravity changelog
Replit and context from all projects
October 6 — Replit announces that it now has context from all of a user’s projects. From a new conversation, users can ask it to find an existing project, add a feature to it or use one project as a reference for another; Replit then reads the relevant files and launches changes as background tasks (background tasks), with links to review the changes. The chosen example goes beyond pure code: applying the color palette of a “Partner Marketing Hub” to two other projects for a coffee brand. The announcement consists of a tweet accompanied by a video, with no blog post, documentation or pricing.
GitHub stacked pull requests become generally available
October 6 — GitHub makes stacked pull requests (stacked pull requests), in public preview since July 30, available on all github.com plans: a large change is split into smaller pull requests, reviewed separately and then merged together. GitHub Enterprise Server will receive them in a future release. According to GitHub, repositories that use stacks merge 9% more code than comparable repositories, and more than two-thirds of the top 1% of repositories use them, with a 5% improvement in time to merge.
The final release reduces friction during merging. When the main branch advances, “Rebase stack” preserves approvals for unchanged code and produces signed replacement commits; a stack passes through the merge queue (merge queue) as a single group, and a stack whose base branch is deleted is retargeted instead of closed. Automatic merging of an entire stack is coming “over the next few weeks.” For command-line use and agents, GitHub CLI’s gh stack extension supports Git worktrees.
🔗 GitHub changelog on stacked pull requests
Multilingual models: Cohere’s Tiny Aya L2-Thinker and TII’s Falcon-OCR-Arabic
Two small models target languages less well served by large models: one reasons in the user’s language, the other reads documents in Arabic.
Cohere’s Tiny Aya L2-Thinker
October 6 — Cohere tackles the language in which models reason: when asked questions in Spanish, Arabic or Swahili, most think in English before answering. Its research model Tiny Aya L2-Thinker (3.35 billion parameters) reasons in the prompt’s language more than 93% of the time, across 60 languages. The recipe comes down to the data mix: roughly 1.7 million English reasoning examples, generated by gpt-oss-120b, lead it to reason in the user’s language only 12.8% of the time; roughly 5,000 translated examples per language, across 44 languages, raise that rate to 86.1%, and multilingual data without reasoning extend this behavior to other languages. Compared with its twin that reasons in English, accuracy drops by at most two to three points on five of the six benchmarks; only PolyMath, which covers competition mathematics, drops more sharply, for lack of a reinforcement learning stage. The model uses fewer than 5,000 reasoning tokens on average. Models and data are published on Hugging Face under the noncommercial CC-BY-NC 4.0 license; the post presents an arXiv paper dated September 9.
🔗 Cohere’s research post on Tiny Aya L2-Thinker
TII’s Falcon-OCR-Arabic
October 6 — TII, the Emirati institute that develops the Falcon models, introduces Falcon-OCR-Arabic on the Hugging Face blog, an Arabic version of its Falcon OCR text recognition model. It retains its 270 million parameters and early fusion architecture, and was adapted through supervised fine-tuning on real and synthetic Arabic documents, followed by reinforcement learning. On a benchmark built by TII itself (11,974 real documents, 15 categories), it ranks second among the 17 models compared, and first on official documents, administrative forms, receipts and invoices.
| Model evaluated by TII | Text accuracy | Table TEDS score |
|---|---|---|
| Gemini 3.5 Flash | 84,34 % | 51,30 % |
| Falcon-OCR-Arabic (270M) | 81,87 % | 59,95 % |
| Claude Opus 5.5 | 79,22 % | 43,98 % |
| GPT Astra | 75,02 % | 51,23 % |
| Chandra OCR 2 (best dedicated OCR) | 61,67 % | 32,63 % |
The post offers only a testing space (playground): no Falcon-OCR-Arabic weights were available on the Hub when we checked, and no license is mentioned.
🔗 TII’s post on Falcon-OCR-Arabic
Hugging Face libraries: Diffusers 0.41.0 integrates Qwen-Image 2.1, Transformers v5.19.0 adds EmbeddingGemma 2
Hugging Face’s two major model libraries release a new version on the same day.
Diffusers 0.41.0 and Qwen-Image 2.1
October 6 — Diffusers 0.41.0, Hugging Face’s library for diffusion models, integrates Qwen-Image 2.1, Alibaba’s model that combines image generation and editing: it can now be used directly to generate, edit, produce images with transparent backgrounds (native RGBA output) and train LoRA adapters, with a visual generation component of 7 billion parameters. The team is also changing its release cadence: like Transformers, Diffusers will now align its minor releases with the arrival of new models, while patch releases remain reserved for fixes. Loading with tensor parallelism allows each GPU to read only its share of the weights, and the release adds LTX-2.5 DFR pipelines and optimizations for Cosmos 3. In return, ONNX support is deprecated in favor of Optimum.
🔗 Diffusers 0.41.0 release notes
Transformers v5.19.0
October 6 — Six days after v5.18.0, Transformers v5.19.0 adds EmbeddingGemma 2, introduced above, with its vectors that can be shortened and its vision and audio encoders that can be disabled at load time. On the distributed training side, expert parallelism gains token distribution (token dispatch), enabled by default for Qwen3 MoE and Mellum: its size no longer has to equal that of tensor parallelism, and Trainer works with this mode. The cache can now be configured layer by layer, which is useful for heterogeneous models, and continuous batch processing (continuous batching) supports XPU devices. The release includes six breaking changes, including returning router logits for all MoE models and phasing out the paged| prefix.
🔗 Transformers v5.19.0 release notes
Voice agents: ElevenLabs launches ElevenAgents Architect in Alpha
October 6 — ElevenLabs integrates an expert called Architect into ElevenAgents, its conversational agent platform, to help teams build and improve their voice or text agents by speaking or writing to it. It knows all ElevenAgents settings (knowledge base, prompts, workflows, voices, guardrails, tools, simulations) and how they interact. Users can give it a question, a failing test, an increase in handoffs to a human or a finding from ElevenAgents Spotlight: it analyzes transcripts, traces the cause, proposes a change and writes the simulations that validate it. It also builds an agent from a simple description. Every change arrives as a versioned, reversible draft, and nothing goes into production without approval. Architect is accessible within the product, as well as from Claude, Claude Code, ChatGPT, Cursor and Grok Bot. It is available now in Alpha, with no pricing or results figures provided.
🔗 ElevenLabs’ post on ElevenAgents Architect
Generative video: FLUX 3 leads Physics-IQ Verified according to Black Forest Labs
October 6 — Black Forest Labs claims first place on Physics-IQ, Google DeepMind’s benchmark measuring video models’ understanding of the physical world: the model is shown the beginning of a filmed real-world experiment and must predict what happens next (fluids, optics, solid mechanics). The reference leaderboard, Physics-IQ Verified, is maintained by Anates Labs. On the video-to-video track, FLUX 3 [large] clearly outperforms NVIDIA’s Cosmos3, at a higher cost per video.
| Model and method (video to video) | Physics-IQ Verified score | Cost per normalized video |
|---|---|---|
| FLUX 3 [large], best of 8 generations | 64,35 % | 16,43 dollars |
| FLUX 3 [large] | 61,11 % | 2,08 dollars |
| Cosmos3 Super (NVIDIA) | 50,80 % | 0,82 dollar |
| Cosmos3 Nano (NVIDIA) | 43,00 % | 0,82 dollar |
On the image-to-video track, FLUX 3 [large] also overtakes Physis-Lang, the method NVIDIA placed first on September 29. FLUX 3 [large] is a proprietary model, and the thread gives neither an availability date nor pricing for this variant.
🔗 @bfl_ai’s thread on X 🔗 Physics-IQ Verified leaderboard
Public evaluations: GeoGuess Bench and RSI Arena
Two evaluations open to the public depart from the usual benchmarks: one has models play GeoGuessr, while the other lets the public vote on models trained by agents.
GeoGuess Bench
October 6 — Hassan, head of developer experience at Together AI, releases GeoGuess Bench, a personal benchmark that has models play GeoGuessr: each sees the same 210 street photos and places a pin scored using the game’s formula, for up to 25 000 points per five-round game. Claude Opus 5.5 takes the lead, just ahead of Claude Fable 5.1; the surprise comes from Muse Glimmer 30B, Meta’s open-weight model, which ranks third, ahead of GPT-6 Astra, at roughly 45 times lower cost per game.
| GeoGuess Bench rank | Evaluated model | Score out of 25 000 | Cost per game |
|---|---|---|---|
| 1 | Claude Opus 5.5 | 23 412 | 0,064 dollar |
| 2 | Claude Fable 5.1 | 23 307 | 0,115 dollar |
| 3 | Muse Glimmer 30B (open) | 22 407 | 0,0049 dollar |
| 4 | GPT-6 Astra | 21 562 | 0,223 dollar |
| 5 | GPT-6.1 Sol | 21 194 | 0,042 dollar |
| 6 | GLM-5.3 Flash (open) | 21 183 | 0,0087 dollar |
GLM-5.3 Flash thus matches GPT-6.1 Sol at roughly five times lower cost. This is the personal project of an employee at Together AI, an inference provider for open models, rather than a company evaluation; no Gemini models are included, and the code is published on GitHub.
🔗 Hassan’s announcement on X 🔗 GeoGuess Bench
RSI Arena
October 6 — Zichen Chen announces a new round of RSI Arena (for recursive self-improvement, recursive self-improvement), this time with Hugging Face. AI agents were given GPUs and a single mission: train better models. They selected the data, wrote the code, and ran the experiments themselves. With the first training round complete, the public joins the loop: the models compete in head-to-head matchups, everyone votes, and the agents use that feedback for new experiments. Hugging Face’s Inference Endpoints serve each model live, and the site lists ten agents, including GPT-6 Astra, Grok 4.7, DeepSeek V4.1 Flash, GLM-5.3, and Muse Spark 1.3; its dashboard showed 286,60 dollars in API spending out of 300 and 755,17 GPU hours out of 1 000. The team promises to publish every agent-trained model on the Hub and, at the end of the experiment, all artifacts: data, code, and each agent’s complete research trajectory.
🔗 Zichen Chen’s announcement on X 🔗 RSI Arena
GPU infrastructure: NVIDIA releases AICR v1.0
October 6 — NVIDIA releases version 1.0 of AI Cluster Runtime (AICR), an open project aiming to end trial-and-error configuration of GPU Kubernetes clusters. An accelerated cluster depends on dozens of components with independently managed versions (kernel, drivers, container runtimes, networking, storage, operators, libraries), and a combination that works in one place may silently fail elsewhere. AICR addresses this with validated, version-pinned recipes, rendered for Helm, Argo CD, Flux, or Helmfile and accompanied by signed validation evidence from the tested hardware. Version 1.0 establishes a compatibility contract: the CLI, REST API, Go SDK, deployment package structure, and artifact schemas become stable interfaces, and any breaking change will require a new major version. NVIDIA claims more than 100 contributors, nearly half from outside the company, and cites integrations at Pulumi Labs and in Mirantis’s k0rdent multicluster manager.
Claude Startups: up to 7 000 dollars in products and credits, and a 45 000-dollar Startup Stack
October 6 — Anthropic expands access to Claude Startups, its program for founders building on Claude, to startups founded less than five years ago or funded within the past two years.
| Program benefit | Value announced by Anthropic |
|---|---|
| One year of Claude Team, up to five Premium seats | 6 000 dollars (five seats at 100 dollars per month) |
| One-time API credit, available upon approval | 1 000 dollars |
| Total Claude products and credits | up to 7 000 dollars |
| Claude Startup Stack (partner offers) | up to 45 000 dollars |
The year of Claude Team applies to companies new to Team, and its actual value depends on the number and type of seats. The Claude Startup Stack, introduced today, brings together discounts and credits from companies building with Claude, including Linear, Lovable, ElevenLabs, Granola, and Hex; its maximum value reflects the combined value of all offers at September 2026 list prices, and not all offers can be combined. The program adds opportunities to speak with Anthropic’s Applied AI team, help publishing a listing on the Claude Marketplace, and access to events.
🔗 Anthropic’s post on Claude Startups 🔗 @claudeai’s thread on the Claude Startup Stack
Briefs
- Claude Projects and local folders — Anthropic’s Dan Fein announces that Claude Projects cloud sessions can connect to an approved folder on the computer, accessing it only when a task requires it; rollout has been gradual since October 5, and the team is almost there (Almost there), according to Boris Cherny. 🔗 source · 🔗 Boris Cherny
- Claude in Slack group messages — Claude can now be added to Slack group messages (group DMs) like any other participant; it responds in a thread that it continues to follow and can use the personal connectors of the person who calls on it, with no details on plans or timing. 🔗 source
- Boris Cherny’s approach to prompting — Boris Cherny has Opus 5.5 build an interactive companion site for an Acquired podcast episode about Home Depot, watercolors included; according to him, there is no secret to prompting: tell the model what you want, how much effort to spend on it, and how to verify the result. 🔗 companion site · 🔗 his approach to prompting
- Atlassian and OpenAI — Under a new agreement, frontier models in the GPT-6 family, including GPT-6 Astra, will power agents on the Atlassian platform and in Rovo; more than 3 000 Atlassian developers already use Codex, and deeper Jira integrations are being explored, with no financial terms disclosed. 🔗 source
- Codex widgets in ChatGPT for iOS — Version 1.2026.267 of ChatGPT for iOS, released on October 2, adds home screen widgets for recent Codex tasks on selected computers, others for remaining usage and its reset time, also available on the lock screen, and a setting that keeps the keyboard open. 🔗 source
- Gemini CLI v0.63.0 and v0.64.0-preview.0 — Stable v0.63.0 reproduces the 15 entries from the September 29 preview exactly, while preview v0.64.0-preview.0 combines the 16 entries from the September 30 to October 3 nightly releases: no new content, and the October 6 nightly release is empty. 🔗 source
- Mistral Python SDK 3.1.0 — Automatically generated from the API, this version adds fourteen evaluation operations in beta under the observability module; the
Runsmethods move underWorkflows.runs, which may break existing code despite being only a minor version change. 🔗 source - Qwen Code v0.25.1-preview.0 — The day after v0.25.0, this preview brings 13 features and 27 fixes, including an experimental Kubernetes runtime, Shell commands and background monitoring for the Managed Agent, and MCP rules that no longer allow a server with the same name. 🔗 source
- SpaceXAI TypeScript SDK 0.2.2 — Released on October 5 without an announcement, this version contains just one change: the
retryBeforeOutputoption also applies to a call tocreate()without a continuous stream (streaming) that expects JSON. 🔗 source - Falcon-Emirati-7B, the Emirati dialect — TII specializes Falcon-H1-Arabic 7B for Emirati Arabic: 84,83 % on Alyah, a benchmark of 1 173 questions, and dialect fidelity of 0,52 versus at best 0,05 for ALLaM, Gemma 3 27B, Jais-2, and Fanar-2, according to a Gemini 3.7 Flash judge; the model is available only on TII’s chat platform. 🔗 source
- Datasets 5.1.0 — Released on October 5, the dataset library now reads Harbor reinforcement learning environments, the Vortex columnar format, and five biological formats (FASTA, FASTQ, GenBank, PDB, mmCIF), and fixes a vulnerability that allowed an archive to write to a neighboring directory. 🔗 source
- TRL v1.14.2 — This patch fixes silently corrupted training: without vLLM, GRPO, RLOO, and Distillation did not stop at the end of the turn for models such as Gemma or Phi-3.5 and learned all subsequent text up to the maximum length; three crashes are also fixed. 🔗 source
- Jev against campaign emails — In a community post, Stephen Solka sorts his political emails with Jev (99 correct answers out of 100 real emails, one false positive), then with a 22,6-million-parameter SetFit classifier that runs on a CPU: 98 % on the pilot evaluation, but 18 out of 23 on difficult cases. 🔗 source
- Open Together on October 16 — vLLM and Hugging Face are organizing Open Together, a gathering of open source developers that will kick off Open Source AI Week on Friday, October 16, in San Francisco; according to Jeff Boudier, 150 people are on the waiting list and capacity has been doubled. 🔗 source · 🔗 Jeff Boudier
- Copilot CLI 1.0.93-2 — This preview adds an enterprise setting,
permissions.limitTo, that restricts network requests to managed domains, highlights GPT-6.1 Sol, GPT-6 Astra and Luna, and Claude 5.5 models in the picker, and now reads user settings only from~/.copilot/settings.json. 🔗 source - AI Scan in the security overview — Organization and enterprise administrators can now see, in GitHub’s security overview (security overview), which repositories have enabled AI analysis of pull requests (AI Scan for pull requests), with two filters and a column in the CSV export. 🔗 source
- Secret detection: Lovable, Pydantic, and Supabase — GitHub’s secret detection (secret scanning) recognizes five new secret types, including the Lovable API key, the Logfire token, and the Pydantic AI Gateway key, as well as two Supabase tokens; Lovable Labs also joins its partner program. 🔗 source
- Devin release notes for October 5 — Mostly refinements: a Terminal tab in the session workspace (opening Files or Terminal wakes Devin), effort levels for Devin Review adjustable through trigger actions, and code analyses (code scans) retrievable by ID through API v3 or Devin’s MCP. 🔗 source
- Delta and its own worktrees — On Delta’s blog, Conrad Irwin explains why the tool replaces Git worktrees: each thread has its worktree recorded in DeltaDB and replicated across people, agents, and machines, multiple agents work on the same branch, and a version-controlled
.agents/preparescript prepares each checkout; the post accompanies no new release. 🔗 source - v0: personal accounts become team spaces — Personal v0 accounts become team workspaces, with conversations, projects, and settings migrated automatically; the documentation nevertheless reserves certain features, such as team templates, for Plus, Business, and Enterprise plans. 🔗 source
- v0 and ChatGPT subscriptions on iOS — ChatGPT subscriptions, accepted by v0 since September 29, can now be used in its iOS app, with no pricing or further details. 🔗 source
- Eleven v4 and Eleven v4 Turbo lead the rankings — ElevenLabs announces that its two speech synthesis models launched on September 28 hold the top two spots on Artificial Analysis’s speech synthesis (text to speech) leaderboard; v4 was already first at launch. 🔗 source
- HeyGen Video in 2K — One week after its launch, HeyGen Video moves to 2K resolution; HeyGen says it ranks first on OpenRouter’s video leaderboard by usage, with no pricing specific to 2K, while the catalog page still lists 0,01 dollar per second through the end of October. 🔗 source
- Nano Banana 2.1 on Pika — Pika adds Nano Banana 2.1, the new version of Google’s Nano Banana model, to its platform and Pika API Club, with better visual quality and greater speed according to Pika; no pricing or credit cost is given. 🔗 source
- DOCA GPUNetIO — NVIDIA makes DOCA GPUNetIO the shared implementation of GPU-driven networking (GDA-KI) for NCCL, NVSHMEM, and NIXL/UCX, available as a full version in the DOCA SDK and a lighter open source project focused on RDMA Verbs. 🔗 source
- CUDA green contexts — A technical post from NVIDIA shows how to reserve SMs for a critical task: on a Blackwell GPU with 148 SMs, a kernel allocated 8 SMs responds in 0,007 ms, compared with 0,140 ms using a priority stream (stream) and 3,727 ms without priority. 🔗 source
- Open models at telecom operators — In a position paper, NVIDIA cites its 2026 telecom survey, in which 89 % of respondents consider open source important, and testimonials from SoftBank, AT&T and Indosat; no new product. 🔗 source
- Perplexity guide to collaboration tools — Perplexity compares ten collaboration tools with AI features (Slack, Loom, Confluence, Airtable, Notion, ClickUp, Asana, Monday.com, Trello, Miro), their AI features and the plans that include them; no new product. 🔗 source
What this means
Decision models are becoming a category in their own right, with their own price war and leaderboard. On the same day, OpenAI sets the price of its Decisions API at $0.10 per million input tokens, and Perplexity lowers its price to $0.02, half what it was five days ago; neither charges for output. These models do not write, they choose or score: users pay for what they read and want them to be fast, with OpenAI promising responses roughly ten times faster than with its Responses API. The Decision Index 0.3 serves as a community arbiter, and its new private half aims to measure skills rather than success on familiar benchmarks. Its verdict already carries weight in vendors’ messaging: Perplexity cites it in its announcement, even though its model card reports a different score from the leaderboard.
AI is moving into office tools rather than the other way around. Claude is entering Google Docs, Sheets and Slides after Excel, PowerPoint and Word, joining Slack group messages, and GPT-6 models will power Atlassian’s Rovo agents. In these integrations, the question is no longer just model quality but control: a mode that requires approval for every change, connectors that follow Google’s sharing permissions, Enterprise controls applied to the add-on. The same logic runs through today’s agent tools, from review decisions that can be undone in Antigravity to versioned drafts in ElevenAgents Architect.
In cybersecurity, access is tiered according to verification. Anthropic’s CVP does not change the model but its guardrails: on CyScenarioBench, the same Opus 5.5 goes from tasks blocked at the very first prompt without verification to 34 tasks completed out of 50 with Red Team Access. Mistral puts forward another argument, that of a model that does not refuse: it claims 82 % on a cyber test that Claude Opus 5.5 and GPT-6 Astra, it says, refuse to take. The two approaches converge on one point: the broadest access to cyber capabilities goes first to verified professionals, whether Mistral partners before the weights are released or members of Anthropic’s program.
Models are also stretching toward both ends of the size spectrum. At the top, Mistral Large 4 and its 1,000 billion parameters, with weights promised for late October; at the bottom, EmbeddingGemma 2, which fits into around 567 MB of RAM on a phone, Tiny Aya L2-Thinker and its 3.35 billion parameters, or Falcon-OCR-Arabic and its 270 million, currently available only as a demo. Many of today’s figures, however, are measured by the vendors themselves: Mistral’s scores, the top spot claimed by Black Forest Labs, TII’s in-house benchmark, Perplexity’s model card, Ironclad’s tasks scored by OpenAI, or the merge improvements announced by GitHub. This is often the only measurement available on announcement day; it would benefit from cross-checking against independent evaluations, which the Mistral Large 4 weights will enable once released.
Sources
- Mistral post on Mistral Large 4
- Mistral Large 4 page in the documentation
- Anthropic post on Claude for Google Workspace
- @claudeai on X, Claude for Google Workspace
- Anthropic post on the Cyber Verification Program
- Cyber Verification Program help page
- Comcast and Booz Allen with Claude Mythos
- EmbeddingGemma 2 on the Google blog
- EmbeddingGemma 2 model card
- @GoogleDeepMind on X, EmbeddingGemma 2
- EmbeddingGemma 2 on Hugging Face
- OpenAI Decisions API guide
- OpenAI API changelog
- @perplexitydevs on X, pplx-decider-v1.1-27b
- pplx-decider-v1.1-27b on Hugging Face
- @multimodalart on X, Decision Index 0.3
- The Decision Index 0.3
- OpenAI post on its collaboration with Ironclad
- @OpenAIDevs on X, usage tiers
- OpenAI API rate limit documentation
- OpenAI help article on the API BAA
- Perplexity API changelog
- Perplexity image_search tool documentation
- GLM 5.3 page on Amazon Bedrock
- @Zai_org on X, GLM-5.3 on Bedrock
- Claude Code 2.1.290
- Claude Code 2.1.291
- Claude Code 2.1.292
- Field guide to cloud sessions on claude.dev
- @ClaudeDevs on X, cloud session bonus credit
- Mistral Vibe v2.26.0 release notes
- Google Antigravity changelog
- @Replit on X, context from all projects
- GitHub changelog on stacked pull requests
- Cohere research post on Tiny Aya L2-Thinker
- TII post on Falcon-OCR-Arabic
- Diffusers 0.41.0 release notes
- Transformers v5.19.0 release notes
- ElevenLabs post on ElevenAgents Architect
- @bfl_ai on X, FLUX 3 and Physics-IQ
- Physics-IQ Verified leaderboard
- @nutlope on X, GeoGuess Bench
- GeoGuess Bench
- @my_cat_can_code on X, RSI Arena
- RSI Arena
- AICR v1.0 on the NVIDIA technical blog
- Anthropic post on Claude Startups
- @claudeai on X, Claude Startup Stack
- Dan Fein on X, Claude Projects and local folder
- Boris Cherny on X, gradual rollout of Claude Projects
- Dan Fein on X, Claude in Slack group messages
- Boris Cherny on X, Acquired companion site
- Boris Cherny on X, how he prompts Claude
- Atlassian and OpenAI expand their partnership
- ChatGPT and Codex changelog, ChatGPT for iOS 1.2026.267
- Gemini CLI v0.63.0
- Mistral Python SDK v3.1.0
- Qwen Code v0.25.1-preview.0
- SpaceXAI TypeScript SDK v0.2.2
- TII post on Falcon-Emirati
- Datasets 5.1.0
- TRL v1.14.2
- Stephen Solka’s post
- @vllm_project on X, Open Together
- Jeff Boudier on X, Open Together waiting list
- Copilot CLI 1.0.93-2
- GitHub changelog on AI Scan enablement status
- GitHub changelog on new secret detectors
- Devin release notes for October 5
- Conrad Irwin’s post on the Delta blog
- v0 changelog, personal accounts converted to team spaces
- @v0 on X, ChatGPT subscription on iOS
- @ElevenLabs on X, Eleven v4 in the lead
- @HeyGenDev on X, HeyGen Video in 2K
- @pika_labs on X, Nano Banana 2.1 on Pika
- DOCA GPUNetIO on the NVIDIA technical blog
- Green contexts on the NVIDIA technical blog
- Open models and telecom operators, NVIDIA blog
- Perplexity guide to collaboration tools