Mistral drops 1T-parameter open-weight multimodal model beating closed frontier models
Mistral just dropped Mistral Large 4, nicknamed Le Chonk, and it’s a big one. We’re talking a 1 trillion parameter model, but here’s the clever part: it only activates 49 billion of those parameters at a time. Think of it like a massive library where you only pull the books you actually need, keeping things fast and efficient.
What can you do with it right now?
- Hit the API today via Mistral Studio at $1.36/M input and $4.18/M output tokens
- Send it images natively, it reads text and visuals together
- Feed it a 1 million token context window, roughly 750,000 words at once
- Use it across 160+ languages
- Run cybersecurity tasks that closed models typically refuse
The open weights drop end of October, meaning you’ll be able to self-host and fine-tune it on your own hardware. Built entirely in Europe on Mistral’s own GPU cluster, it’s the strongest open-weights model from the US or Europe on aggregated benchmarks.
Claude integrates directly into Google Docs, Sheets, and Slides for live editing
Claude just moved into your Google files. No more copy-pasting between tabs.
Two ways this works. Install the add-on from the Google Workspace Marketplace, and Claude appears as a sidebar inside your open Doc, Sheet, or Slide. Or flip on the connectors inside Claude, paste a Google file link, and the file opens right next to your chat.
Here is what you can actually do:
- Docs: fix sentences, rewrite sections, or restructure content without breaking your formatting
- Sheets: write formulas, build pivot tables, run Python on a data range, and write results back into the sheet
- Slides: build new slides using your existing layouts or flag things that look off
You stay in control. Bigger edits show up as suggestion cards you accept or dismiss. Your Google sharing permissions carry over, so nothing leaks outside what you already shared.
Available now in beta on all paid Claude plans via the Google Workspace Marketplace.
OpenAI drops new math results from a frontier model advised by top mathematicians
OpenAI just dropped 722 math manuscripts from an unreleased internal model onto GitHub. These aren’t textbook exercises. We’re talking about real open research problems, many of which have stumped mathematicians for decades.
Here’s how it happened. The model was given roughly 4,000 problems to attempt. Each result used about three hours of compute. The outputs that cleared a significance bar got packaged into 372 families of related papers, complete with Lean proofs, meaning a computer can actually verify the math is correct.
What’s in the repo? Things like:
- Progress on the irrationality exponent of pi
- A proof of the Hodge Conjecture for a specific class of geometric objects
- Results on the Riemann zeta function’s zero-free region
The release follows pressure from the math community. 25 Fields Medal winners signed an open letter warning AI labs were racing to solve famous problems without proper oversight. OpenAI consulted an independent advisory group of top mathematicians to shape how this release was handled. The repo is Apache-2.0 licensed and publicly accessible.
Mistral built a trillion-parameter model and is opening the weights this month
Mistral Large 4 landed yesterday in API preview, and the internal nickname tells you the shape of it: le Chonk. It is a 1 trillion-parameter natively multimodal mixture of experts with 49 billion active parameters, a hybrid instruct-and-reasoning model trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s European datacenters, across more than 160 languages. The weights ship by the end of October.
The benchmark Mistral leads with is a security one. ML4 solves 93% of the 40 challenges in Cybench, a set drawn from capture-the-flag competitions, and places in the top five of the AA Cyber Index. On the agentic side it posts 61.7% on DeepSWE v1.1, 28.3% on Terminal-Bench 4, and 59.9% on AutomationBench, which runs 657 business workflows across Gmail, Google Sheets, Slack and Salesforce. Mistral also claims it beats GPT-6 Astra on visual grounding, at 42% on Dense 200.
The number that actually moves anything is the price. ML4 runs $1.36 per million input tokens and $4.18 per million output. Astra, per OpenAI’s own GPT-6 build guide, is $10 and $50. You do not have to believe ML4 is Astra’s equal to notice that an open-weight model a third of Astra’s price, with a cyber score that good, changes what a self-hosted security or compliance stack costs to run.
Worth watching what happens when the weights actually drop. A 1T model with a 93% Cybench score, downloadable by anyone, is the exact artifact that Anthropic spent yesterday building an access-tier system around. Mistral is skipping the tiers.
OpenAI put 372 families of new math results on GitHub
OpenAI published a dump of new mathematical results yesterday, produced by an unreleased internal frontier model and posted to github.com/openai/math. The repo holds 722 manuscripts organized into 372 families, with Lean formalizations, 10 summaries of the model’s reasoning, and the compute accounting. The model was posed roughly 4,000 problems, and the average result cost about three hours of ChatGPT Pro thinking.
That compute figure is the interesting disclosure. Three Pro-hours per result, times 372 families, is a research program you could price out on a credit card rather than a supercomputer allocation. OpenAI frames the release around its work with the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and promises workshops on how to handle machine-produced results.
The field is not uniformly thrilled. Scientific American’s writeup ran under the headline that OpenAI “unleashes hundreds more math results upon a field already in shock,” which is a fair read of the problem: none of this is peer reviewed, “new” is not the same as “important,” and 372 families out of 4,000 attempts is a hit rate, not a discovery rate. Nothing here says which results matter.
It is also the second lab in a week to do this. On Monday we covered Meta publishing six math papers written with Muse Spark, five of them closing previously open problems. Meta shipped six papers with authors attached. OpenAI shipped 722 manuscripts with a compute estimate. The second approach scales, and that is exactly what mathematicians are nervous about.
Anthropic opened its cyber models wider, and published the vulnerability count
Anthropic folded Project Glasswing and its original Cyber Verification Program into one tiered program yesterday. Verified security professionals now pick a lane: Defense Access for SOC work, incident response and malware analysis, Red Team Access which adds authorized penetration testing, and Specialized Access for testing safety-critical systems like flight controls, power grids and telecom networks.
The receipts are the point. Program partners turned up 129,000 verified vulnerabilities between April and July 2026, and Anthropic’s own work found 5,500 more from April through October, with over 33,000 rated critical or high severity. Partners including Booz Allen and Comcast told Anthropic that Claude Mythos raised their rate of vulnerability finding “by months or even years.”
The tier math shows what verification actually buys you. On CyScenarioBench, Defense Access gets blocked on 46 of 50 tasks. Red Team Access gets blocked on zero. That is the whole product in two numbers: the classifiers are not tuned down a little for cleared users, they come off. Which makes the verification step, and not the model, the thing worth auditing.
Anthropic gives startups a free year of Claude Team
Anthropic widened its Claude for Startups program at SF Tech Week. Now, approved startups get a free year of Claude Team for up to five premium seats, plus $1,000 in API credits.
The Claude Startup Stack adds partner offers worth up to $45,000 from tools such as Hex, ClickHouse, Granola, Linear, Lovable and ElevenLabs. Startups backed by an investor in Anthropic’s partner network can get up to $100,000 more in credits. The free year is for companies new to Team.
However, the credits expire after six months and they work only on Anthropic’s own API, not Bedrock or Vertex AI. So, plan your build around that. If you want more details or want to check if you qualify for the program, visit the Claude for Startups page.
You can also join the Anthropic team at Claude Founder House from SF Tech Week through October 8. Sessions cover raising money, building agents and running an AI-native company.
Anthropic opens its most powerful models to more security teams
Anthropic merged two programs into one expanded Cyber Verification Program. Qualifying security teams can now work with fewer safeguards on Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1.
The expansion follows Project Glasswing, whose partners found at least 129,000 verified software vulnerabilities between April and July. More than 33,000 were rated critical or high severity. The figures cover only some partners, so Anthropic expects the true count to be at least five times higher.
Security teams, critical infrastructure operators, open-source maintainers and researchers with a record of reported vulnerabilities can apply through the application page.
To see how Comcast and Booz Allen use Claude Mythos to find exploit chains and secure their codebases, read the analysis here.
Meta and Sierra propose a standard for how personal AI agents deal with businesses
Sierra and Meta published the Personal Agent Protocol, an open standard for how personal AI agents sign in to businesses. It uses OAuth, so a customer can give an agent read or write access to an account. Each business chooses how agents reach it: through its website, its APIs or its own agent. Genesys, Instinct, Rocket, Shopify, Stripe and Walmart are founding partners. A first spec is due later in October.
They already have a rival. Visa’s Trusted Agent Protocol launched in June, and Stripe and Shopify have joined that one too. Amazon, which has blocked Meta’s agents, isn’t a partner. OpenAI and Anthropic aren’t listed as well.
Right now, every agent and every business sorts this out on its own, which Sierra’s Bret Taylor calls chaos. For brand and CX leaders, the protocol would show when an agent is acting for a customer and let the company decide what that agent can do. Visa’s rival protocol already counts Stripe and Shopify as members, so businesses may end up supporting both standards.
AI Scenario Planning Assistant
When to use this?
Use this when an AI trend or business change could affect your plans. Explore possible outcomes, spot early warning signs, and decide what to do next.
You are an AI Scenario Planning Assistant designed for enterprise AI leaders and mid-level executives. Help users anticipate possible business outcomes, evaluate uncertainty, and make better-informed strategic and operational decisions. Your job is to turn a business decision, emerging trend, AI initiative, or uncertain situation into a set of plausible scenarios with clear implications and actionable next steps. ### How you should work When a user shares a business challenge or decision: 1. **Understand the context:** Identify the business objective, decision to be made, time horizon, key constraints, and uncertainties. Ask targeted follow-up questions only when essential information is missing. 2. **Identify key drivers:** Surface the internal and external factors that could materially affect the outcome, including AI adoption, technology costs, regulation, customer behavior, competition, workforce readiness, and organizational capabilities. 3. **Build plausible scenarios:** Develop 3–4 distinct scenarios, such as a baseline, an upside case, a downside case, and a disruptive or unexpected case. Make each scenario internally consistent and grounded in the user's context. 4. **Assess business implications:** For each scenario, explain the potential impact on revenue, costs, productivity, customers, workforce, operations, and AI investments, wherever relevant. 5. **Identify signals and triggers:** Highlight the early indicators that could suggest a scenario is unfolding, what to monitor, and what developments should prompt a change in strategy. 6. **Recommend practical actions:** Separate no-regret moves that make sense across scenarios from actions that should be taken only if specific conditions emerge. Identify opportunities, risks, dependencies, and trade-offs. ### Output format Present your analysis in a concise, executive-ready format: * **Decision or question:** What the user is trying to resolve. * **Key uncertainties:** The factors that could change the outcome. * **Scenario overview:** A table comparing each scenario, its assumptions, potential business impact, and implications. * **Early warning indicators:** What to monitor and why. * **Recommended actions:** Immediate steps, contingency plans, and decisions to revisit. * **What to watch next:** The most important questions or signals over the next 30, 90, or 180 days, depending on the user's time horizon. ### Important guidelines * Clearly distinguish facts, user-provided information, assumptions, and hypothetical scenarios. * Do not present speculative outcomes as predictions or assign probabilities unless there is a defensible evidence base. * Where quantitative estimates are useful, explain the assumptions and calculations. Never invent company data, financial figures, or market forecasts. * Keep scenarios meaningfully different, not just minor variations of the same outcome. * Tailor the analysis to the user's industry, company size, strategic priorities, resources, and risk tolerance. * Prioritize clarity and actionable decisions over exhaustive analysis or jargon. * Help users understand the trade-offs and make their own decisions rather than prescribing a single outcome. Your ultimate goal is to help enterprise leaders prepare for multiple plausible futures instead of relying on a single forecast or reacting to events after they happen.
Paramount Seals $110B Warner Deal
Paramount Skydance completed its roughly $110 billion acquisition of Warner Bros. Discovery, bringing HBO, CNN, CBS and major Hollywood franchises under one roof. Led by David Ellison, the merged company plans to combine streaming services and target at least $6 billion in cost savings, raising concerns over layoffs and media consolidation.
Lambda Seeks Up to $4B Ahead of IPO
Nvidia-backed AI cloud provider Lambda is seeking up to $4 billion at a $14.5 billion pre-money valuation, with Coatue and Blackstone leading the round ahead of a planned 2027 IPO. The fundraising comes as Lambda’s reported contract backlog has surged to $50 billion, highlighting both strong demand for GPU capacity and the enormous costs of expanding AI infrastructure.
Google Signs 20-Year Nuclear Deal
Google is updating six US nuclear plant sites and signing a 20-year power purchase agreement as it scrambles for electricity to feed its data centers. Big Tech is now underwriting baseload power directly, turning energy into a core AI constraint.
Anduril Wins Navy Contract Worth Up to $2.9B
Defense startup Anduril secured a US Navy contract worth up to $2.9 billion to manufacture submarine components, alongside a $3.7 billion investment in a new Maryland production facility. The deal comes shortly after co-founder Palmer Luckey joined a Pentagon weapons advisory initiative, underscoring the growing role of defense technology firms in US military manufacturing.
Vinci Raises $250M for Simulation
Engineering simulation platform Vinci4D closed a $250 million Series B at a $1.5 billion valuation, co-led by Advent, Temasek and Xora, with AMD Ventures participating. Hardware teams are buying more virtual prototyping as physical testing gets slower and costlier.
Claude Connects to Google Workspace
Anthropic added a sidebar integration that lets Claude work inside Google Docs, Sheets and Slides, going straight after Gemini’s home turf. The move pushes AI assistants out of the chat box and into the documents where teams actually work.
OpenAI Claims 100+ Math Solutions
OpenAI says its latest internal model has solved over 100 long-standing math problems, with results yet to be published. The claims have sparked criticism from mathematicians over attribution, independent verification and bypassing traditional academic publishing.
Xanadu and GlobalFoundries Partner on Quantum Chips
Photonic quantum computing company Xanadu signed a multi-year partnership with GlobalFoundries to develop and manufacture key quantum computing components using established 300mm semiconductor production lines. The collaboration will initially focus on photon detectors and low-loss photonic circuits, aiming to move quantum hardware manufacturing beyond laboratory-scale production.
Xbox Secures Exclusive GTA 6 Streaming
Xbox CEO told staff Microsoft is preparing something around Grand Theft Auto VI that no other platform holder is doing, reportedly locking up exclusive streaming rights. It signals the next console battle will be fought over cloud access, not just hardware.
Open Cosmos Raises 300M Euros for Satellites
UK-based Open Cosmos closed a 300 million euro Series C and plans to launch 192 low-Earth-orbit satellites for its ConnectedCosmos constellation, pitching sovereign-grade space services. Europe is funding its own orbital infrastructure to reduce dependence on US providers.
Tesla Expands EV Home Backup Power
Tesla is expanding vehicle-to-home backup power to newer Model 3 and Model Y vehicles, allowing owners with a Powerwall 3 to use their car batteries during electricity outages. The system could provide several days of additional backup power, although these models do not yet support sending electricity back to the grid. The move brings Tesla’s vehicles further into its home energy ecosystem.
Meta Upgrades Ray-Ban Smart Glasses
Meta is rolling out audio improvements and turn-by-turn walking navigation to its Ray-Ban smart glasses. The updates keep refining the everyday-wearable bet that glasses, not headsets, become the mainstream spatial-computing form factor.
Deliveroo Plans UK Robot Delivery Rollout
Deliveroo and Coco Robotics announced a partnership to introduce autonomous pavement delivery robots across the UK, beginning in London’s Canary Wharf in October. Additional launches are planned for Milton Keynes, Leeds, Stockton-on-Tees and Nottingham over the coming months. The rollout marks Deliveroo’s first use of autonomous delivery vehicles in the UK, with robots initially handling short-distance orders alongside human couriers.
Hadrian Raises $40M for AI Security
Dutch offensive-security startup Hadrian raised $40 million to expand internationally, using AI agents to find the weaknesses attackers would exploit. Automated, continuous attack-surface discovery is becoming a funded category of its own.