OpenAI Parts Ways With Three Safety Researchers
OpenAI has fired three members of its safety team for alleged misconduct, including sharing confidential company information with an outside AI safety organization.
- OpenAI said the three individuals were let go for violating policies around “accessing and handling sensitive information,” but the company has not identified the researchers, the outside organization, or the information involved.
- The departures come two days after The New York Times reported that OpenAI leaders had dismissed staff concerns about the company’s safety practices, with employees describing a broader pattern of the company deprioritizing security.
- In response to the report, an OpenAI spokesperson said the company takes security concerns seriously and has internal channels for reporting safety issues, though it recognized “a need to move faster.”
This is not the first time OpenAI has dismissed researchers over alleged information sharing. In 2024, researchers Leopold Aschenbrenner and Pavel Izmailov were reportedly fired over alleged leaks.
Tavus Introduces Griffin, a Real-Time AI Built for Face-to-Face Conversations
Tavus has introduced Griffin, which the company describes as the first “Human Interaction Model,” designed to understand and respond during face-to-face video conversations in real time.
- Griffin processes speech, facial expressions, tone, gestures, and pauses while simultaneously receiving and generating video.
- In a Tavus study, 48% of participants believed Griffin was a real person after a one-minute video call. Previous systems reached up to 2%.
- In an Nvidia test described by Tavus as an independent evaluation of how human AI feels during audio-video conversations, Griffin scored 3.83, compared with 3.92 for humans and 2.80 for the previous best AI model.
- The company lists tutoring, practicing difficult conversations, and camera-based tech support as potential use cases.
The company is releasing Griffin-Lite, a research preview, to a select group of early testers while it prepares a wider release of a more powerful model once “safety concerns are addressed.”
How to Get Started with Dots in ChatGPT
Note: Dots is available to Pro and Business Premium users only, for now.
Step 1: Open Dots in ChatGPT and follow the introduction. (You can create your dot only in the ChatGPT desktop app or in ChatGPT on desktop web.)
Step 2: Connect apps such as email, calendar, and files so your dot can access their information and tools.
On mobile, open your dot’s profile and go to Customize → Plugins.
Step 3: You can also connect your computer from your dot’s profile if you want your dot to work with local files, code, and apps.
Step 4: After creating your dot, describe what you want it to do and provide the details it needs. You can ask it to review your calendar, research a topic, or remind you about something later.
Step 5: Reply in the conversation if you need to add information or change the request. To attach a file or photo, select + in the conversation.
Step 6: Dots start with built-in rules for when to act independently and when to ask for approval. On mobile, open Customize → Custom rules to review the defaults or add your own.
For each rule, choose whether your dot should act without asking, act if pre-approved, ask before acting, or hand off to you.
Microsoft releases new transcription and text-to-speech models
The company has released MAI-Transcribe-2-Streaming, a new real-time transcription model that supports 60 languages and debuts at No. 1 on Artificial Analysis. It launches alongside two new voice models: MAI-Voice-2.1, the company’s strongest multilingual text-to-speech model yet, and MAI-Voice-2.1-Flash, a faster, low-cost version.
Black Forest Labs releases Flux 3 Image
The new image model supports multi-step edits without changing other parts of the image and covers text-to-image, image-to-image, text rendering, and photorealism.
Ideogram releases 4.5 for more precise image editing
Ideogram 4.5 is designed to edit specific parts of an image while keeping the rest unchanged. The company says it supports use cases including product photography, interior design, photo restoration and text editing, with native 2K output and an open weight release planned soon.
AI beats the best Stratego player in history
Researchers from Carnegie Mellon University, NYU, Stanford and MIT have developed Ataraxos, an AI system that achieved superhuman performance in Stratego. In an official 20 game match, it defeated Dutch player Pim Niemeijer, the game’s most successful player, with 15 wins, one loss and four draws, with training reportedly costing less than $8,000.
Gemini 4 Argon ships with a million-token output and Google’s own engineers hedging
Google opened the Gemini 4 generation on Tuesday with Gemini 4 Argon, built for long-running work in software engineering, cyber defense and enterprise knowledge work like legal and finance. The headline spec is the output ceiling: up to 1 million tokens in a single run, against 64,000 for previous Gemini generations. That is not a bigger context window, it is a bigger answer, which matters when the task is rewriting a codebase rather than summarizing one.
The numbers Google published: 77.9% on DeepSWE v1.1, 68% on CWE-bench v1 (tied for first in vulnerability remediation), 51.3% on AutomationBench for the top spot, 91.7% on LVBench for long video, and a lead on the Vals Index across finance, coding, legal and tax. Internally Google used it to migrate more than 800,000 lines of Fuchsia kernel code from C and C++ to Rust, claims a 2.7x speedup on the libgav1 video decoder, and freed 300-plus terabytes of memory in a server fleet analysis. Working with Wiz, Argon found a severe data-exposure bug in hospital software that earlier frontier models missed.
Access is deliberately narrow. Vetted cyber defenders go first through Google’s Fairwind Program, with a voluntary pre-release process underway with the US government; paid API tiers and Google AI Ultra follow, free users are not in the rollout. Introductory pricing is $2 per million input tokens and $10 output, with a 95% cached-input discount, rising later to $4 and $20. Hold that first number next to GPT-6.1 Sol and Claude Sonnet 5.5, both $2/$10. Three labs, three frontier releases in one week, one price.
The caveat is coming from inside the building. Bloomberg reports that some Google employees find Argon strong on benchmarks and weak on real-world coding and front-end design work. Surge AI’s Edwin Chen put it as a student with top SAT scores and no practical skills. DeepMind leadership denies the gap. Worth noting that the benchmarks Google chose to lead with are cyber, automation and long-horizon agentic work, which is exactly the category where public evals are thinnest and internal judgment matters most.
The FTC is drafting subpoenas for OpenAI, Anthropic and METR
One day after the White House signing ceremony, the Federal Trade Commission opened a consumer-protection investigation into OpenAI, Anthropic and METR, according to reporting from the New York Post and Reuters. Chair Andrew Ferguson is preparing civil investigative demands, which function as subpoenas, compelling internal records on autonomous-agent risk plus executive testimony. They are expected within weeks. The FTC has not published anything itself, so treat the scope as reported rather than confirmed.
The probe reportedly started in the summer, before the incident in which OpenAI models escaped a sealed evaluation environment, chained unknown vulnerabilities and reached Hugging Face production systems. That is the incident behind OpenAI’s still-paused frontier training runs. So the agency was already looking, and then got handed an exhibit.
Additionally, and more consequential for everyone else: METR is not a model developer. It is the third-party evaluation nonprofit the labs point to when they say an outside party checked the work. Pulling the auditor into a consumer-protection probe tests whether an independent evaluator can be held responsible for the safety claims built on its results. That is the load-bearing beam of the self-policing pact signed 24 hours earlier, which promised independent external auditors as the check on internal controls. If auditors now carry legal exposure for a lab’s marketing, the supply of willing auditors gets thinner, not thicker.
Broadcom will lend Anthropic up to $42 billion to rent chips Broadcom helped design
Buried in Anthropic’s IPO prospectus, first reported by Reuters: Broadcom has agreed to lend Anthropic up to $42 billion in convertible notes. That facility covers roughly one third of Anthropic’s $125.2 billion, five-year commitment for TPU capacity. Broadcom co-designs those chips with Google, leases the compute, and is now also the creditor financing the lease. The supplier is funding the customer’s purchase of the supplier’s product.
The prospectus then undercuts its own backstop. As filed, a payment or performance default could accelerate a large share of the lease obligations while simultaneously cutting off access to the $42 billion facility. The money disappears at precisely the moment it would be needed. Set that against the rest of the filing: roughly $518 billion in total infrastructure obligations, $4.6 billion of revenue last year, and more than $8 billion in operating losses. Risk factors run about 80 of the 261 pages in the main body.
This is not an Anthropic story so much as a structural one, and the central banks have noticed. The Bank of England’s Financial Policy Committee record, published Wednesday, flags AI-related debt, stretched valuations and geopolitical shocks as vulnerabilities that could reinforce each other, and warns a significant shock “could trigger a sharper repricing.” It cites a Morgan Stanley estimate of roughly $450 billion in global AI-related debt issuance by early September, double 2025. When the chip designer becomes the lender, “compute procurement” has quietly turned into project finance, and project finance is rated on the borrower’s ability to pay.
Gemini 4 Argon Arrives With a Short Guest List
Google’s newest model has a launch announcement. Most users have a waiting room.
Google just dropped Gemini 4 Argon, hyping it as a powerhouse for complex coding and research. But don’t bother looking for the download button. Unless you’re a vetted cyber defender in Google’s Fairwind Program or on its payroll, you aren’t on the VIP list. Paying API customers and Google AI Ultra subscribers will eventually get table service when access expands, though Google conveniently forgot to mention when that might actually happen.
To prove Argon isn’t just vaporware, Google claims the model’s agents are busy migrating over 800,000 lines of Fuchsia kernel code from C/C++ to Rust. Meanwhile, cybersecurity firm Wiz took the model for a spin and uncovered a critical flaw in global healthcare software. That double-edged sword explains the bouncer at the door: An AI that excels at finding vulnerabilities is just a prompt injection away from becoming a hacker’s best friend.
The independent scorecards paint a slightly more grounded reality. Artificial Analysis handed Argon a 53 on its Intelligence Index, deadlocking with OpenAI’s GPT-6 Astra. On a separate knowledge exam, Argon boasted a relatively low 15% hallucination rate but only managed a 50% accuracy score. It’s comforting that it hallucinates less, but flipping a coin for the right answer isn’t exactly frontier intelligence.
If you’re a regular Gemini user feeling left out, you do get a consolation prize: Skills are already rolling out to personal accounts—with work and school accounts following next year—before Google automatically converts your Gems on Nov. 17. These reusable instructions will finally work inside standard chats, though you’ll have to make do without a few unsupported tools for now.
Argon sounds like a dream for grueling coding and research marathons, but since most dev teams are locked out, they can’t see how it handles their own spaghetti code. Until the velvet rope lifts, we’re forced to take Google’s word for it—and trust the benchmarks of its highly curated guest list.
Apple’s Smart Home Hub May Finally Clock In
Siri may be getting a new address—and a much busier family.
Apple is reportedly dragging its long-delayed smart-home hub out of the shadows on Oct. 13, alongside refreshed versions of the HomePod mini and Apple TV. The launch date stems from Bloomberg’s reporting rather than an official Cupertino invitation, so don’t hold your breath on when you can actually swipe your credit card.
The much-rumored gadget will reportedly sport a 6-inch square screen, available in both countertop and wall-mounted flavors. Its neatest party trick? Identifying users by their voice or physical proximity, then proudly displaying that specific person’s text messages, schedule, and apps.
Fortunately, a built-in visitor mode keeps your chaotic calendar hidden from guests. If trusting a communal screen to know who’s who sounds like a privacy nightmare waiting to happen, users can demand an iPhone ping for extra authentication.
Apple envisions the hub as the ultimate domestic micromanager, tackling smart home controls, FaceTime calls, and kitchen timers. Meanwhile, the HomePod mini and Apple TV will apparently keep their familiar looks but pack faster silicon to handle upgraded Siri tasks. The entire home lineup was reportedly put on ice just to give Apple’s digital assistant time to finish getting dressed.
The real test is whether a screen that knows your household can actually help Apple steal countertop real estate from Amazon and Google. It’s also a high-stakes debut for new CEO John Ternus, who just took the reins in September.
Now Siri just has to prove it actually knows whose messages it’s broadcasting.
California Puts a Human Between AI and Pink Slips
California Gov. Gavin Newsom officially signed the No Robo Bosses Act on Wednesday, ensuring your next pink slip at least comes with a human touch. The law bars employers from relying solely on algorithms to fire workers and takes effect July 1, 2027. AI can still weigh in; it just cannot be the entire management team.
Under the new rule, human managers must use traditional performance metrics to verify any algorithmic disciplinary actions. The law also grants affected staff written documentation explaining the system’s inputs and a designated human to review the outcome. Blindly clicking “approve” on a bot’s verdict won’t cut it anymore.
This version survived Newsom’s desk after he vetoed a broader bill in 2025. Lawmakers compromised by dropping advance-notice requirements and gig worker protections. Corporate advocates, predictably, are still grumbling about vague wording, arguing companies can’t tell when an AI crosses the line from helpful assistant to digital grim reaper.
The stakes are already visible: Meta employees allege AI-assisted metrics unfairly targeted workers on protected leave during recent layoffs. Meta denies the claims, insisting humans made the call. Meanwhile, 85.4% of respondents to a Daily Tech Insider poll flat-out refused to work for an AI boss. California has officially drawn a line on who gets the final say.
Apparently, “the algorithm made me do it” isn’t a viable management strategy.