ai-powered-markdown-translatorArticle translated from French to English using gpt-6.1-sol.
On Friday, October 2, Meta publishes six mathematics papers developed by researchers with its Muse Spark model; according to Meta, five of them answer research questions that had remained open. Anthropic commits 100 million dollars to train 10,000 deployment engineers through the Claude Frontier Academy and releases Claude Code 2.1.288, while NVIDIA announces a 64 GB DGX Spark starting at 4,999 dollars. Two announcements from the previous evening round out the day: Perplexity’s Decisions API, which answers with probabilities using a model released with open weights, and Google’s satellite carrying its first TPUs into orbit.
Meta publishes six mathematics papers written with Muse Spark, Ai2 open-sources AstaBrief, Google’s AI forecasts flu
October 2 — Meta publishes six mathematics papers developed by researchers with Muse Spark, its in-house model, on its research blog. According to Meta, five of them answer previously open research questions, though the post does not specify which one is the exception. Over the past few months, the mathematicians worked with Muse Spark 1.1 and 1.2 in Thinking mode, directly in the chat interface at meta.ai, without custom research scaffolding (custom research scaffold).
Following gold-medal-level performance from our AI models across five competitions in mathematics, physics, and chemistry, we asked a harder question: can AI contribute when a problem is genuinely open and without an existing solution path? — @AIatMeta on X
The stated protocol is strict: one team of mathematicians guides the research, a second group reviews the work, and each paper identifies the passages written primarily by researchers or by AI. The model’s contribution varies considerably from one paper to another, from supporting calculations to writing entire sections:
| Mathematical field | Result announced by Meta | Muse Spark’s contribution according to Meta |
|---|---|---|
| Probability | Sharp threshold for fitting an ellipsoid through random Gaussian points in high dimensions | Proof strategies |
| Differential equations | Finite-time blow-up of negative-energy radial solutions to the mass-critical biharmonic nonlinear Schrödinger equation, a question open since 2015 | Calculations, testing arguments, revising the proof |
| Group theory | A semi-abelian group is not necessarily monomial: a counterexample of order 384 to a conjecture by M. Kida (2024) | Search program written in GAP, which found the counterexample |
| Optimization | Exact rule identifying when a cycle relaxation captures the original problem | Probabilistic reformulation, counterexample, proof strategy |
| Arithmetic physics | The two-point function of p-adic string theory equals a height function, over a class of curves much broader than the Tate curve | Candidate proofs, three technical sections written |
| Non-associative algebra | A counterexample of dimension 3 to a conjecture on solvable evolution algebras | Counterexample and alternative characterizations |
Meta acknowledges that three of these problems have also been solved elsewhere: three papers published in August on the Gaussian threshold, a counterexample to Kida’s conjecture reported on September 16 by the AI agent Nilradical, and counterexamples on evolution algebras by Hu and Wen. According to Meta, its own results were obtained independently, using different approaches. The announcement follows OpenAI’s, which said on September 21 that an internal model had solved more than 100 open problems; Meta emphasizes the use of a consumer model as-is and transparency about AI’s contribution, with all proofs still checked, corrected, and rewritten by mathematicians.
🔗 Solving Open Research Problems Together (Meta AI Research)
AstaBrief 8B: Ai2 open-sources a model that writes scientific reports with citations
October 2 — Ai2 open-sources AstaBrief 8B, a model that turns a research question and excerpts from the literature into a scientific report with citations. It now powers the Fast mode of the “Generate a report” feature in Asta, Ai2’s agentic platform for science, alongside the Thinking mode powered by Claude. The recipe remains simple: Qwen3-8B, supervised fine-tuning (SFT) on 47,000 examples produced by the ScholarQA pipeline, followed by direct preference optimization (DPO) on around 6,000 pairs, without reinforcement learning. The weights, licensed under Apache-2.0 according to the model card, and the training data are publicly available: a lab can run it on its own hardware.
The model writes its report in a single pass. Across Asta’s entire pipeline, a report takes an average of 51.1 seconds in Fast mode versus 178.5 in Thinking mode, making it around 3.5 times faster. In a small human study (14 questions, 3 researchers), DR Tulu wins on overall preference, but two out of three researchers prefer AstaBrief for citation accuracy. Ai2 cautions that most of the training and evaluation dates back to 2025 and has not been repeated against current models.
🔗 Open-sourcing AstaBrief (Ai2)
Flu: a model developed with Google’s AI ranks first out of 39 in the CDC’s FluSight evaluation
September 30 — Google announces that a flu forecasting model developed with its AI was the best of the 2025-2026 season in FluSight, the program run by the US Centers for Disease Control and Prevention (CDC) that aggregates hospital admission forecasts every week from October to May. The CDC’s end-of-season report confirms this: of the 39 models retained from the 53 submitted by 34 teams, the top individual model is Google_SAI-FluEns, with a relative WIS of 0.56 (an error score: the lower it is, the better the forecast), ahead of three models at 0.58, including OHT_JHU-nbxd. The FluSight ensemble, which the CDC uses for its public communications, ranks 7th.
These forecasts were developed with Empirical Research Assistance (ERA), a Google AI tool that generates optimization algorithms for various scientific fields and is available to trusted testers. The associated research paper describes an autonomous system that writes, evaluates, and improves its own forecasting code through an LLM-guided tree search; it has also produced models for COVID-19 and respiratory syncytial virus (RSV).
🔗 Google Research blog post · CDC FluSight 2025-2026 report
Perplexity launches the Decisions API and open-sources pplx-decider-v1-27b, llama.cpp serves decision models
October 1 — Perplexity launches a new API, the Decisions API, and releases the model that powers it, pplx-decider-v1-27b, under the Apache 2.0 license. The announcement came in the evening through @perplexitydevs, the company’s developer account. Instead of writing an answer, this decision model (decision model) returns a probability distribution over a fixed set of answers: it reads text, JSON, or images, but does not write an answer, generate code, or explain its reasoning.
Three question types are supported: noul returns the probability of a yes, choice chooses from 1 to 255 options with a probability for each and a confidence score, and score places the content on a scale of up to 10 levels. A request can ask up to 128 questions about the same content and must remain below 262,144 input tokens, including the questions. Pricing is 0.04 dollars per million input tokens, output is free, and there is no per-request charge; each organization can send 10 requests per second, regardless of its plan. Perplexity targets classification, routing, and rubric-based scoring: its reference example (cookbook) sorts support tickets, with the Decisions API making the initial decision and the Agent API handling escalations.
The model is fine-tuned from Qwen3.8-27B; its weights, available on Hugging Face, require a CUDA GPU capable of holding around 49 GiB, plus working memory. Its model card compares it with Jev, another decision model, and with its base model across eleven benchmarks, with pplx-decider-v1-27b’s results measured through Perplexity’s API:
| Benchmark (accuracy) | Jev score | Qwen3.8-27B score (base model) | pplx-decider-v1-27b score |
|---|---|---|---|
| WinoGrande | 90,70 % | 73,10 % | 83,30 % |
| FinancialPhraseBank | 76,98 % | 75,68 % | 84,18 % |
| RAGTruth | 77,27 % | 61,53 % | 88,80 % |
| JudgeBench | 78,57 % | 68,86 % | 78,29 % |
| BBH | 94,27 % | 72,80 % | 82,80 % |
| JevBench public hard | 73,27 % | 72,28 % | 70,30 % |
| TabFact | 89,80 % | 78,60 % | 90,60 % |
| ContractNLI | 77,45 % | 80,78 % | 80,78 % |
| Circa | 84,60 % | 87,00 % | 89,20 % |
| Belebele | 95,00 % | 93,20 % | 94,00 % |
| TruthfulQA binary | 92,00 % | 82,80 % | 85,40 % |
| Overall (according to the model card) | 84,51 % | 74,76 % | 85,71 % |
On the Overall row, highlighted in the announcement, pplx-decider-v1-27b leads Jev, 85.71% versus 84.51%, but the model card does not specify how this row is calculated. According to the published table, Jev remains ahead on six of the eleven benchmarks, including BBH and WinoGrande, and on JevBench public hard, Perplexity’s model even falls behind its base model.
🔗 Announcement from @perplexitydevs · Decisions API documentation · pplx-decider-v1-27b on Hugging Face
llama.cpp serves decision models with the Jev-compatible /v1/systemone API
October 2 — The llama.cpp server can now serve decision models through a new endpoint, /v1/systemone, introduced in PR no. 29818, merged on October 2. The ggml-org blog post, written by Xuan-Son Nguyen and Victor Mustar, describes how it works: you send a state (text, JSON, screenshot) and typed questions, choice, score (from 2 to 10 levels), or noul (yes or no), and the model returns a probability for each option in a single pass. The API adopts the System One format introduced by Jev, TypeSafe’s model: an existing client only needs to change its base URL. A router mode loads multiple models on demand on the same server.
| Supported decision model | Model size | Base model | Median time per question |
|---|---|---|---|
| Julia-1 (more than 50 languages) | 144M | mmBERT-small | 3 ms |
| Laya | 421M | ModernBERT-large | 5 ms |
| Kev-4B | 4B | Qwen3.5-4B-Base | 12 ms |
| lev | 4B | Qwen3.5-4B | 36 ms |
| OpenJev (also reads images) | 27B | Qwen3.8-27B | 43 ms |
The timings are measured on an NVIDIA RTX PRO 6000. The next announced model is Clef, which Cloudflare introduced on October 1 with an API compatible with Jev’s; ggml-org created the GGUF repositories for Clef and Clef-flash on Hugging Face on October 2.
🔗 New in llama.cpp: Decision Models
Anthropic launches Claude Frontier Academy and commits 100 million dollars to train 10 000 engineers
October 2 — Anthropic is launching Claude Frontier Academy, a training program backed by a commitment of 100 million dollars. The goal is to train 10 000 frontier deployed engineers (Frontier Deployed Engineers, FDE) by the end of 2027. The first cohorts bring together engineers from Accenture, Bain, Capgemini, Commonwealth Bank of Australia, Deloitte, McKinsey, Morgan Stanley and Novo Nordisk.
The first program, the Frontier Deployed Engineer Residency, follows the medical training model. According to the program page, it begins with an intensive four-day module: three days in person building a Claude system for a simulated company, followed by a graded assessment that awards the Claude Resident Engineer badge. This is followed by a twelve-week residency implementing a Claude deployment within the participant’s own organization, then a final assessment for the Claude Frontier Deployed Engineer badge; the first badges are expected in early 2027.
Access remains restricted: participation is by nomination, early access is reserved for selected customers and partners, and cohorts take place in San Francisco, New York and London. The Academy extends the Claude Partner Network (46 000 companies, more than 175 000 certifications), launched in March with 100 million dollars; the post does not say whether the new funding is additional.
🔗 Anthropic’s announcement · Program page
Agents: Claude Code 2.1.288, GLM 5.3 in Cursor, Replit, DeepSeek Harness and the SWE-sweep benchmark
Claude Code 2.1.288 removes the limit on background commands in interactive sessions
October 2 — Claude Code 2.1.288 arrives with 89 entries, including 64 fixes. The most visible change concerns commands running in the background: the time limit introduced in 2.1.285 (30 minutes by default, 2 hours maximum) now applies only to unattended sessions (-p, Agent SDK, CI, cloud). In the terminal, desktop application and VS Code, a background command can once again run indefinitely.
A prompt accidentally cleared with Ctrl+C can be recovered with the up arrow, /code-review accepts --max-findings to report more or fewer issues, claude project purge becomes claude purge, and mods, launched the previous day, receive $.ui.selection(), which returns the last text selected in full screen. When the API times out mid-response, non-interactive sessions and subagents resume from the partial response instead of failing. On security, a dangerous rm slipped into a bash -c script no longer goes through without confirmation in bypassPermissions mode, and a PreToolUse or PermissionRequest hook whose matching fails now blocks the call instead of being ignored. Finally, the client-side auto mode classifier now ignores an ANTHROPIC_DEFAULT_SONNET_MODEL variable that specifies Sonnet 5.5 or Opus 5.5, and uses Claude Sonnet 5 instead.
Six minutes after the release, @ClaudeDevs highlighted You should know, a built-in mod that launches a secondary agent tasked with displaying information users should not miss above the prompt. It was already among the six built-in mods presented on October 1 and remains disabled by default.
We’re adding a new plugin to Claude Code: You should Know. It scans Claude’s output for important information you might miss to help keep you in the loop. Enable it with: /plugin enable cc-plugin-you-should-know@builtin — @ClaudeDevs on X
🔗 Claude Code 2.1.288 release notes
GLM 5.3 and GLM 5.3 Flash arrive in Cursor
October 1 — Cursor adds GLM 5.3 and GLM 5.3 Flash, Z.ai’s open models, to its model selector, and claims that GLM 5.3 Max ranks first among open models on its own benchmark.
GLM 5.3 and GLM 5.3 Flash are now available in Cursor! GLM 5.3 Max is the best-scoring open-weight model on CursorBench 4.0. — @cursor_ai on X
The attached chart plots models by their CursorBench 4.0 score and average cost per task, but labels only one point: GLM 5.3 marked “(max)”, at 42.6% for 5.05 dollars per task. According to this chart, GLM 5.3 remains below Claude Opus 5.5, Claude Sonnet 5.5, Fable 5.1 and Grok 4.7; no other open model is shown, and GLM 5.3 Flash does not appear. Cursor has published no pricing, post or methodology to support its ranking.
Replit makes GPT-6.1 Sol and Claude Sonnet 5.5 available in Agent, and Jev in its AI Integrations
October 2 — Replit’s changelog adds two recent models to the choices available in Replit Agent: Claude Sonnet 5.5 in Power mode and GPT-6.1 Sol in Max mode. The same entry makes Jev, the decision model previously mentioned alongside pplx-decider-v1-27b, available in Replit AI Integrations: an application can classify content, route requests or score leads through fast, structured decisions, without managing an API key. The documentation specifies that Jev is served under the identifier jev-latest through OpenRouter. Settings now open as a full page, and Enterprise accounts can set company-wide rules, with exceptions for individual teams (Workspace). The entry provides no pricing.
🔗 Replit’s October 2 changelog
DeepSeek Harness comes to macOS and Windows
September 30 — The @DeepSeekHarness account announces a desktop version of DeepSeek Harness, DeepSeek’s open source agent harness, for macOS and Windows. On October 2, the main @deepseek_ai account shares the announcement and specifies that Linux users should use the npm package @deepseek-ai/dsh. Installers can be downloaded from DeepSeek’s website, for macOS on ARM chips and Windows x64.
The official page, marked PREVIEW, presents Harness as an open source public preview available worldwide. It is built on the Cordis framework’s “everything is a plugin” (everything is a plugin) architecture and covers office work (files, data, documents, presentations), coding, research with source citations and background tasks, with a Creator mode that writes new plugins through conversation. The new development concerns distribution: the September 29 preview already integrated the dsh command into the desktop application, and no stable version has yet been published on GitHub.
🔗 Announcement from @deepseek_ai · DeepSeek Harness page
SWE-sweep: a benchmark where agents must find bugs to fix on their own
October 1 — Researchers from Meta Superintelligence Labs, together with Harvard, the University of Washington and Stanford, release SWE-sweep, a benchmark that asks coding agents to discover bugs in a real repository themselves and fix as many as possible, without hints about their nature or location. The dataset covers 100 repositories and 4 068 bugs; the code, licensed under MIT, uses the Harbor framework, and Ofir Press and John Yang are among the authors. No announcement yet accompanies the repository, created on October 1.
The leaderboard, updated on September 24 and measured using mini-SWE-agent, shows the scale of the task:
| Leaderboard rank | Model evaluated | Bugs resolved | Total cost |
|---|---|---|---|
| 1 | Sol 5.6 (xhigh), OpenAI | 4,7 % | 7 230 dollars |
| 2 | Luna 5.6 (xhigh), OpenAI | 2,5 % | 224 dollars |
| 3 | Terra 5.6 (xhigh), OpenAI | 1,5 % | 357 dollars |
| 4 | Luna 5.6 (high), OpenAI | 1,4 % | 28 dollars |
| 5 | Opus 5 (xhigh), Anthropic | 1,3 % | 5 363 dollars |
Computing and hardware: first TPUs in orbit, DGX Spark 64 Go, DeepGEMM and DeepEP on Ascend 950
Project Suncatcher: Google’s prototype satellite carries its first TPUs into orbit
October 1 — A week after announcing its first orbital test, Google launched the prototype satellite for Project Suncatcher, a long-term research project (moonshot) exploring whether space could one day host large-scale machine learning infrastructure. Built with Planet, the satellite reached orbit aboard Transporter-18, a SpaceX rideshare mission; according to Google’s Travis Beals, the team has established contact and the satellite is operating as expected.
Over the coming weeks, Google will measure how its TPUs handle the stresses of spaceflight and extremes of radiation and temperature; on the ground, Trillium TPUs had already withstood a radiation dose greater than that of a five-year mission. The peer-reviewed paper detailing the research appears in the journal Joule. Sharing the launch on October 2, @GoogleAI reiterates the project’s rationale: in low Earth orbit, sunlight is nearly constant, and satellites can produce up to 8 times more solar energy than on Earth. The announced next step is two satellites in 2027 to test laser links.
🔗 The Project Suncatcher prototype is in orbit · The paper published in Joule · The post shared by @GoogleAI
DGX Spark 64 Go: NVIDIA adds a configuration going on sale on October 23
October 2 — NVIDIA adds a configuration with 64 Go of unified memory to its DGX Spark lineup, alongside the 128 Go model. It will be sold exclusively by six partner manufacturers, Acer, ASUS, Dell, Gigabyte, HP and MSI, starting Friday, October 23, at a starting price of 4 999 dollars; the post does not give the current price of the 128 Go model. The foundation remains the same: a GB10 Grace Blackwell chip, DGX OS and the NVIDIA AI software stack, for models with up to 100 billion parameters running on the machine.
The main selling point is clustering: two units connected by a QSFP cable through their ConnectX-7 network cards pool their memory.
| Technical specification | One DGX Spark 64 Go | Two DGX Spark 64 Go in a cluster |
|---|---|---|
| Unified memory | 64 Go | 128 Go pooled |
| Advertised model size | Up to 100 billion parameters | Up to 200 billion parameters |
| Performance on Qwen 3.8 27B (NVIDIA test) | Baseline | Up to 1,7 times |
The NVIDIA Sync application handles this with Cluster Assistant, which detects the units and configures the network. NVIDIA Sync Model Launcher, announced for the end of the month, will download and launch Qwen3.8 27B on a machine or cluster, and configure OpenCode to use it.
🔗 NVIDIA’s post on DGX Spark 64 Go
DeepSeek ports DeepGEMM and DeepEP to Huawei’s Ascend 950 NPUs
September 29 and 30 — DeepSeek published two ports of its computing libraries to Huawei’s Ascend NPUs on GitHub, without an announcement. DeepGEMM-Ascend, whose README dates the first release to September 30, reproduces the DeepGEMM API exactly: matrix multiplications in BF16, FP8 and FP4, MQA logits and MegaMoE, under the MIT license, for the Ascend 950 series. DeepEP-Ascend ports the communication library for expert parallelism (expert parallelism) in MoE models, with all-to-all distribution (dispatch) and recombination (combine) operations.
On Ascend 950DT, DeepEP-Ascend measures dispatch throughput of 373 to 375 Go/s with 8 ranks and 313 to 320 Go/s with 128 ranks; up to 32 ranks, it reaches approximately 90 to 95% of the physical bandwidth limit. These figures were obtained with a proof-of-concept hardware development kit (HDK) supplied to DeepSeek and a manual configuration, neither publicly distributed; the README points to the commercial version, which Huawei expects to make available around October 15.
🔗 DeepGEMM-Ascend on GitHub · DeepEP-Ascend on GitHub
Voice and audio: Suno Speech in beta, ElevenLabs certified FedRAMP 20x Class A
Suno Speech generates spoken voice and background music in a single track
October 1 — Suno launches Speech in beta, a feature that creates spoken audio over original background music, directly in Suno: users enter an idea, poem or text, then describe the desired voice and musical style. Suno presents Speech as the first audio model to generate voice and music together in a single coherent track. Tested for a month with a small group of users, the feature is now available to everyone following an application update.
The post, written by Chief Product Officer Jack Brody, mainly cites personal uses: messages from friends turned into dramatic readings, voice notes set to music, meditations or bedtime stories. Suno cautions that the beta remains imperfect: a British accent can drift toward Australian, and dramatic pauses can be very pronounced. No pricing, plans or supported languages are specified.
ElevenLabs certified FedRAMP 20x Class A for ElevenAgents and its voice APIs
October 1 — ElevenLabs obtains FedRAMP 20x Class A certification, the framework through which US federal agencies determine whether a cloud product can be used safely with government data. It covers ElevenAgents and the Text to Speech and Speech to Text APIs, provided they operate in Zero Retention Mode with data residency in the United States. It expands the ElevenLabs for Government offering, and the company is now listed in the FedRAMP Marketplace. According to FedRAMP, cited by ElevenLabs, Class A is sufficient for most non-sensitive agency uses; ElevenLabs cites benefits applications, permits, taxes and hearing transcription.
ElevenLabs supports the announcement with existing public-sector deployments:
| Public-sector deployment cited | Figure reported by ElevenLabs |
|---|---|
| National employment hotline (Ukraine) | 25 % handled by conversational agents |
| Benefits and employment hotlines (Czechia) | 85 % of approximately 5 000 daily calls resolved |
| City of Midland | 7 000 fewer missed calls per month (projection) |
ChatGPT opens Finances to Free and Go users in the United States
October 2 — Finances, ChatGPT’s space for personal finance, is expanding to Free and Go plans in the United States, on the web, iOS and Android. The feature was already available to Plus and Pro subscribers: launched as a preview exclusively for Pro subscribers in May, it added Experian credit score tracking for Plus and Pro subscribers on September 21. It now covers Free, Go, Plus and Pro plans, still exclusively in the United States.
The approach remains the same: users connect their bank and investment accounts through Plaid, and their credit report through Experian, with the two connections available together or separately. The Finances page brings together spending, bills, subscriptions, net worth, investments and credit scores. In conversations, ChatGPT categorizes spending, identifies subscriptions, compares the current month with recent trends and, for taxes, organizes information and highlights options to explore, such as deductions.
🔗 ChatGPT release notes · Finances in ChatGPT
SpaceXAI releases its official TypeScript SDK as an experimental version
October 2 — SpaceXAI released the first public version of its official TypeScript SDK without an announcement. Version 0.1.0 was released at 16:57 UTC, followed at 18:25 UTC by version 0.2.0, which renames the client class xAI to SpaceXAI: a breaking change that aligns the code with the company’s new name. The SDK already covers the Responses API (streaming, compaction and multi-turn conversations), the platform’s built-in tools (web search and X search, code execution, remote MCP servers), image and video generation and editing, as well as the Files, Batch and Voice APIs.
| SDK component | Observed detail |
|---|---|
| npm package | @xai-official/sdk |
| License | Apache-2.0 |
| Requirements | Node.js 22.13 or newer, ESM project, no runtime dependencies |
| Status | Experimental, interfaces subject to change before 1.0 |
By default, the SDK refuses to run in a browser or Worker to avoid exposing the secret key on the client side. TypeScript developers now have an equivalent to the official Python SDK, whose version 1.20.0 was released on September 24.
🔗 SpaceXAI TypeScript SDK CHANGELOG · The @xai-official/sdk package on npm
Briefs
- GPT-6.1 Sol in v0 — On September 29, the model’s release date, v0, Vercel’s app generator, made GPT-6.1 Sol available, a few hours after enabling connections through a ChatGPT subscription; the announcement provides neither pricing nor metrics specific to v0. 🔗 source
- New agent dashboard in Grok Build — Announced on October 1, it displays all agents on one screen with
/dashboard, launches tasks in parallel and places an agent in its own git worktree with Ctrl+W; Grok Build received its first dashboard on June 15. 🔗 source - Gemini CLI nightly — The October 2 v0.64.0-nightly contains just eight fixes: Ctrl+C once again cancels an ongoing operation, and the
~/.gemini/state.jsonstate is written atomically, with a.bakbackup restored in case of corruption; the stable and preview versions are unchanged. 🔗 source - Models removed from GitHub Copilot — As announced on September 3, Gemini 3.5 Flash, Gemini 3.6 Flash, Kimi K2.7 Code and Claude Opus 4.7 are leaving all Copilot experiences; GitHub suggests Gemini 3.8 Flash, Kimi K3 and now Claude Opus 5.5, instead of Claude Opus 5. 🔗 source
- Testing a SKILL.md — In a community post, Golda Manuel proposes a method, without published measurements, for checking whether a skill is discovered, loaded and useful to coding agents: placement where Claude Code or Codex expects it, activation at three request levels, and comparison with and without the skill across eight metrics. 🔗 source
- Manus details Video Editor and Game Dev — Two October 1 posts explore Manus 2.0 in more depth: Video Editor, in Manus Studio, places each video element on its own track (125 clips edited into a 10 min 50 s film with 195 cuts), and Game Dev adjusts a game live; neither pricing nor a date is provided. 🔗 Video Editor · 🔗 Game Dev
- ServiceNow’s AutoSynthData — ServiceNow’s CoreAI team generates verified training tasks from a model’s failures: on EnterpriseOps Gym, Gemma-4-26B-A4B-it gains 7.2 Pass@1 points in the Hybrid domain and rises from 18.77% to 27.18% in ITSM; neither code nor data has been published. 🔗 source
- POCKET-Darwin-180B — VIDRAFT releases a 111 GB 4-bit GGUF version (compared with 360 GB) of Darwin-180B-RSI, a 180-billion-parameter MoE: 18.4 to 21 tokens per second on CPU alone, 4.17 on a laptop with an 8 GB RTX 5060, and 87.65% on MMLU-Pro, matching the original, according to the authors. 🔗 source
- GPT-6 model guide — OpenAI publishes a guide for startups: GPT-6 Astra for the hardest reasoning tasks, GPT-6.1 Sol for complex coding, research and computer use, and GPT-6 Luna for focused tasks at scale; cached tokens cost up to 95% less depending on the model. 🔗 source
- ChatGPT scans multiple pages into a single PDF using the camera on iOS — rolling out, October 1.
- Chatham Financial and Codex — The capital markets consultancy built a transaction validation app with Codex: according to initial measurements, verification drops from about 30 minutes to under 4; its Onyx platform combines GPT-5.6 Sol and Terra, GPT-5.4 and GPT-4.1. 🔗 source
- The Den and ChatGPT Work — According to an OpenAI case study dated October 1, the leadership team of this Denver club for parents saves 10 to 15 hours a week with ChatGPT Work, connected to Gmail, Slack and Google Drive; a grant application takes 2 hours instead of 3 days. 🔗 source
- GPT-6 Astra Ultrafast on Blackwell — NVIDIA stated on October 1 that the GPT-6 Astra Ultrafast tier, launched by OpenAI on September 29, runs on its Blackwell GPUs, up to 8 times faster than Astra Standard mode, and that OpenAI optimizes its inference software with its own models; no new figures were provided. 🔗 source
- RL Rollouts at CoreWeave — Introduced on September 30 alongside CoreWeave Forge and shared by NVIDIA on October 2, this service built on NVIDIA Dynamo hot-loads new checkpoints during reinforcement learning post-training; on Nemotron 3.5 Lightning, model reload latency improves by a factor of 15. 🔗 source
- Eleven v4 in ElevenReader — ElevenLabs’ most expressive voice model, launched on September 28, arrives in its reading app: articles, ebooks and PDFs in the chosen voice, in more than 90 languages, on iOS, Android and the web. 🔗 source
- Midjourney alpha changelog — Dated October 1, it mainly brings together fixes, including prompt settings remembered per folder and style references grouped into a single chip, ahead of new collaboration tools announced for the following week. 🔗 source
- Grok 4.7 at half price in Ramp Router — Until October 6, the Ramp Router LLM gateway offers Grok 4.7 at a 50% discount, an offer shared by SpaceXAI; the model is priced at 2 dollars for input and 6 dollars for output per million tokens on SpaceXAI’s API. 🔗 source
- Confidential comments on security advisories — GitHub allows comments on a repository security advisory without the reporter seeing them: only people with write access can read these comments, and views are logged; a REST API for comments enters public preview. 🔗 source · 🔗 REST API
- GitHub Advisory Database GraphQL API — The SecurityAdvisory object gains five fields, including the CVE identifier and NVD publication date, and the securityAdvisories query gains two filters, by severity and withdrawn status: there is no longer a need to fall back to the REST API. 🔗 source
- End of the macOS 14 image in GitHub Actions — Announced on October 1, its retirement will take place on November 2, preceded by eight scheduled outages in October, from 14:00 to midnight UTC; workflows must move to macos-15 or macos-latest, which points to macos-26. 🔗 source
- Enterprise AI maturity — Perplexity publishes a guide describing four maturity stages, summarizing frameworks from Gartner, Deloitte and Google Cloud, and offering a self-assessment framework; according to the 2026 McKinsey survey it cites, 80% of respondents see personal productivity gains, but only 37% see a contribution to EBIT. 🔗 source
What it means
AI applied to science is increasingly judged by criteria beyond those set by the people announcing it. Meta does more than list results: each paper identifies which passages AI wrote, a second group of mathematicians reviews the work, and the post acknowledges that three of the problems have also been solved elsewhere. The flu forecasting model developed with Google’s AI derives its value from an evaluation conducted by the CDC over an entire season, against 38 other models. Ai2 publishes AstaBrief’s weights and data, but warns that its comparisons date from 2025: an open model can be checked, while an evaluation grows outdated.
Decision models are becoming a category of their own in a matter of days. Following Jev, Liquid AI’s d1 and Cloudflare’s Clef, Perplexity offers both an API billed per input token and its model’s weights; llama.cpp serves these models locally using the same System One format, and Replit adds Jev to its integrations without requiring an API key. For developers, the benefit is concrete: a probability for each option, obtained in a single pass, instead of a text response that then needs to be parsed. Comparisons call for caution: the model card’s Overall row places pplx-decider-v1-27b ahead of Jev, but the published breakdown gives Jev the advantage on six of eleven benchmarks.
Coding agents are advancing faster than their actual autonomy. Left alone with a repository on SWE-sweep, the best model finds and fixes just 4.7% of bugs, at a cost of more than 7,000 dollars. Claude Code 2.1.288 devotes most of its 89 entries to fixes, including recovery after an API timeout and several tightened permission rules, and highlights a mod that monitors the main agent’s work. Anthropic is tackling the other bottleneck, human skills, with 100 million dollars to train engineers capable of taking a deployment through to production. Open models, meanwhile, are earning their place in tools, with GLM 5.3 in Cursor and DeepSeek Harness on the desktop.
Finally, computing hardware is diversifying in three directions. Google is testing TPUs in orbit that could one day benefit from up to 8 times more solar energy than on Earth, NVIDIA is adding a 64 GB configuration to its DGX Spark that can be paired to reach 200 billion parameters, and DeepSeek is making its low-level libraries usable, with the same API, on Huawei’s Ascend chips. Each, in its own way, expands the locations and hardware on which models can run.
Sources
- Solving Open Research Problems Together (Meta AI Research)
- @AIatMeta’s thread on the six papers
- Open-sourcing AstaBrief (Ai2)
- Google Research post on flu forecasting
- CDC FluSight 2025-2026 report
- Decisions API announcement by @perplexitydevs
- Decisions API documentation
- pplx-decider-v1-27b on Hugging Face
- New in llama.cpp: Decision Models (ggml-org)
- Claude Frontier Academy (Anthropic)
- Claude Frontier Academy program page
- Claude Code 2.1.288 release notes
- You should know highlighted by @ClaudeDevs
- GLM 5.3 in Cursor, announcement by @cursor_ai
- Replit changelog for October 2
- DeepSeek Harness shared by @deepseek_ai
- DeepSeek Harness page
- SWE-sweep
- The Project Suncatcher prototype is in orbit (Google)
- Project Suncatcher paper in Joule
- Launch shared by @GoogleAI
- DGX Spark 64 GB (NVIDIA blog)
- DeepGEMM-Ascend on GitHub
- DeepEP-Ascend on GitHub
- Introducing Speech (Suno)
- ElevenLabs certified FedRAMP 20x Class A
- ChatGPT release notes
- Finances in ChatGPT (OpenAI Help Center)
- SpaceXAI TypeScript SDK CHANGELOG
- The @xai-official/sdk package on npm
- GPT-6.1 Sol in v0 (@v0)
- Grok Build agent dashboard (@grok)
- Gemini CLI October 2 nightly v0.64.0
- Deprecated GitHub Copilot models
- Does your SKILL.md help coding agents?
- Manus Video Editor
- Manus Game Dev
- AutoSynthData (ServiceNow)
- POCKET-Darwin-180B (VIDRAFT)
- A model guide for the GPT-6 family (OpenAI)
- Chatham Financial and OpenAI
- The Den and ChatGPT Work
- GPT-6 Astra Ultrafast on Blackwell (NVIDIA blog)
- Introducing CoreWeave Forge
- Eleven v4 in ElevenReader (@ElevenLabs)
- Midjourney alpha changelog for October 1
- Grok 4.7 in Ramp Router (@SpaceXAI)
- Confidential comments on security advisories (GitHub)
- REST API for security advisory comments (GitHub)
- New SecurityAdvisory GraphQL API fields (GitHub)
- Retirement of the macOS 14 image in GitHub Actions
- Perplexity’s guide to AI maturity