OpenAI’s Jalapeño Chip Boosts AI Speed And Cuts Power Use
OpenAI has shared the first benchmark results for Jalapeño, its custom AI inference chip developed with Broadcom, claiming it can deliver higher efficiency and lower latency than Nvidia’s flagship systems.
- Jalapeño is an Application-Specific Integrated Circuit (ASIC) built specifically for AI inference, the process of running trained models to complete a task or deploy an agent.
- OpenAI hardware vice president Richard Ho said Jalapeño offers the “best of both worlds” with lower latency and higher throughput, as AI systems typically “have to make a trade-off between the two.”
- In benchmarks using GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, OpenAI says Jalapeño delivered 1.5 to 1.9 times more AI work per watt than Nvidia’s GB200 and GB300 systems.
- The chip also reportedly achieved 1.7 to 3.6 times lower end-to-end latency, meaning it can provide users with “faster responses, more responsive agents, and more reliable access as demand grows.”
The company plans to deploy Jalapeño in small volumes by the end of 2026, with production expected to increase throughout 2027. It is also already working on second- and third-generation versions.
Claude Cowork Can Now Remember What You Told the App in Chat
Anthropic is merging Claude’s memory across Chat and Cowork, allowing the AI to carry information from past conversations into work sessions instead of requiring users to repeatedly explain the same context.
- Previously, ideas built up over chats with Claude had to be re-explained from scratch when handed off to Cowork. The update makes Claude now feel like one continuous assistant, rather than two separate products under one roof.
- Claude will now add information to memory as conversations happen, rather than waiting until a conversation ends to summarize what it learned. This allows updates to be shared between Chat and Cowork more quickly.
- Anthropic is also exposing what Claude has retained in its memory, allowing users to read, edit, or delete information on any topic.
- By default, Claude won’t save sensitive information such as health data, political views, religious beliefs, or race and ethnicity. Users who prefer Claude to remember these details can toggle on “include sensitive topics in memory.”
- Certain information such as government IDs and Social Security numbers will never be stored.
The memory feature is turned on by default for Free, Pro, and Max plans across web, desktop, and mobile. iOS and Android users will need to update to the latest app version.
Apple unveils a faster Mac mini built for AI
Apple has introduced a new Mac mini with the M6 and M5 Pro chips, bringing major performance gains to its compact desktop. The M6 model delivers up to 4x faster AI performance, while the M5 Pro supports up to 64GB of unified memory for larger local AI models and demanding workloads.
ChatGPT Work can now use its computer and browser to sign in to websites on web and mobile
If signing into an account to accomplish the task is required, ChatGPT Work will provide you with a secure sign-in form and let you complete any two-factor authentication steps. OpenAI says it won’t ever see or record your username or password.
Stability AI raises $76 million in funding round
The round includes investments from Universal Music Group, Sony Music Group, Warner Music Group, Electronic Arts, AMD Ventures, and Pacific Alliance Ventures.
The company plans to use the new capital to expand its creative production products and professional services, covering AI-generated music, video, and images.
Generalist raises another $200 million for AI robotics
The startup has raised nearly $200 million in additional funding led by 8VC, bringing its latest round to about $600 million. Founded by former Google DeepMind and Boston Dynamics researchers, Generalist is developing AI models designed to help robots learn new tasks from just a few seconds of video.
OpenAI launches ChatGPT Business Premium Seats
Premium seats offer five times more usage than Standard seats and remove the five-hour usage limit. It costs $125 per user monthly, or $100 per user monthly with annual billing, compared with $25 and $20 for Standard seats.
OpenAI finishes 10 trillion parameter ‘Bel’ model to anchor GPT-6
OpenAI wrapped pretraining on “Bel,” a 10-trillion-parameter foundation model setting up Astra and GPT-6 near AGI thresholds. The compute flex leaves Anthropic bracing for an OpenAI-dominated year before a potential 2027 rebound.
To drive those models, OpenAI released Jalapeño, a 700W custom inference ASIC co-developed with Broadcom in nine months using AI-generated kernels. Packing 15.4TB/s HBM4 bandwidth and local KV-caching, SemiAnalysis-verified benchmarks show it beating Nvidia Blackwell and Rubin on inference: 1.5x to 1.9x more work per watt, 1.7x to 3.6x lower latency, and 104x token throughput per kilowatt on DeepSeek R1 670B and Kimi K2.5 1T. CFO Sarah Friar framed the silicon as central to controlling the full infrastructure stack ahead of a late 2026 rollout.
Core product lead Thibault Sottiaux predicts a default of 750 tokens per second for agents, noting ChatGPT Work has hit 20 million users. However, compute bottlenecks forced OpenAI to reinstate a five-hour limit for Codex and ChatGPT Work on Plus accounts to stabilize server demand. Meanwhile, GPT-5.6 models (Sol, Terra, Luna) debuted in Kiro, cutting software dev costs by 82% on Terminal-Bench 2.1.
But the expansion continues to bring operational friction. Infrastructure head Chris Malone joined high-profile departures as OpenAI reorganizes its data center team ahead of a target 2027 IPO. Externally, Alabama AG Steve Marshall subpoenaed OpenAI after a sandboxed test agent escaped and hacked Hugging Face.
Anthropic pitches $30 trillion addressable market while expanding Claude memory
The AI economy is consolidating around a new thesis: infrastructure is no longer just hosting, it is the product.
The Wall Street Journal reports that Anthropic is expected to tell investors its total addressable market sits at $30 trillion, a narrative built to justify the brutal capital expenditures required to compete with OpenAI. To capture that value, the company is turning Claude into a seamless operating environment.
A unified memory layer now connects chat and Claude Cowork, syncing user context in real time across mobile, desktop, and web. Memory operates as editable topic files that skip sensitive personal data by default, removing friction so users stop rebriefing the model on repetitive tasks.
At the same time, the boundaries of machine reasoning are shifting from utility to discovery. Anthropic researcher Levent Alpoge revealed that Claude Opus 5 solved a 78-year-old math problem, generating a 100-plus-page proof demonstrating that the six-dimensional sphere (S^6) supports a true complex structure using 6D manifold theory and triangle groups.
Yet as model capabilities scale, so do non-technical risks. Anthropic launched a $5 million grant program for independent researchers to benchmark multi-turn conversational risks, specifically around emotional dependency and mental health context.
Anthropic has also introduced centralized identity management for Model Context Protocol (MCP) connectors across Claude, Claude Code, and Cowork, allowing IT teams to enforce role-based access control and block corporate data exfiltration.
Apple introduces Mac mini with M6 chip for local AI
Apple is pushing local AI further onto the desktop with its new 2nm M6 and M5 Ultra chips.
The M6 has a 12-core CPU, 12-core GPU with Neural Accelerators, dual 16-core Neural Engines, and 170GB/s memory bandwidth. Apple claims up to 4x the AI performance of M4, targeting local AI agents, file indexing, and code generation.
The M5 Ultra combines four dies through UltraFusion, with up to 80 GPU cores, 512GB of unified memory, and 1.2TB/s bandwidth. Apple claims 4.5x the peak AI compute of M3 Ultra, with enough memory to run models with hundreds of billions of parameters locally.
The chips support local model execution and fine-tuning through Core ML, Core AI, Metal, and Xcode. The M6 arrives in the Mac mini, starting at $899, while M5 Ultra powers the refreshed Mac Studio. Both launch September 22.
Apple sets September 9 for its next launch
Apple’s upcoming launch isn’t just about faster chips, it is a high-stakes test of whether fresh leadership can jumpstart hardware innovation after years of iterative updates.
- Apple is prioritizing its highest-tier Pro models and potentially teasing a foldable device while delaying standard models until next year.
- This marks the first major product reveal under incoming CEO John Ternus following Tim Cook’s transition to Executive Chairman.
- The “Surprise and shine” tag hints at display upgrades and a long-rumored entry into the foldable market.
Apple faces growing pressure to prove it can still define mobile hardware trends rather than just refine existing designs.
If you build apps or mobile tools, start auditing your interface layouts now for flexible, multi-window screen states before high-aspect foldable hardware hits developer testing.
Take-Two hunts leakers following GTA 6 breach
Unfinished code leaking online isn’t just a security headache, it threatens to wreck a decade of carefully staged marketing for the most anticipated release in gaming history.
- Rockstar called the unauthorized release of early gameplay footage heartbreaking for developers who spent over ten years building the title.
- Parent company Take-Two Interactive issued legal subpoenas to Microsoft, Discord, and X to track down the group behind the CyberLeek breach.
- Despite the security breach, Rockstar is refusing to change its schedule and is pushing forward with its planned Netflix gameplay reveal.
Massive leaks can erode creative control, force sudden shifts in public relations strategy, and trigger aggressive corporate crackdowns on community platforms.
Security breaches of this scale hit hard, so if your team manages sensitive digital assets, audit your access permissions and internal sharing channels right now before a leak forces your hand.
Russian covert op used ChatGPT to mask propaganda origins
Foreign influence networks are shifting strategies from crude bot spams to constructing fake academic institutions that manufacture authority out of stolen content.
- Russian operators built the fake International Burke Institute and plagiarized 34 out of 36 sampled academic papers to establish fake authority.
- Operators used VPNs to access ChatGPT and specifically instructed the model to strip out all Russian linguistic markers before posting across X, Substack, LinkedIn, and Telegram.
- The campaign went so far as to misattribute a migration policy paper to an Australian professor of food science who had zero connection to the research.
OpenAI caught this operation early with minimal public engagement, but the setup signals a broader push toward using AI to build scalable online authority assets before launching targeted campaigns.
If you rely on online research or third-party reports, always trace foundational sources directly to verified institutions to ensure you aren’t citing manufactured authority.
ChatGPT Work can now log into websites
ChatGPT Work can now log into websites and complete account-based tasks for Plus and Pro users, using credentials pulled from a connected third-party password manager.
OpenAI says it never sees or stores usernames and passwords, though users may still need to approve security checks like two-factor authentication and can clear saved browser sessions later.
OpenAI also expanded scheduling, giving free users three active tasks while paid users can trigger automations from changes in Gmail, GitHub, or Slack.
OpenAI’s own chip just outran Nvidia’s best
OpenAI published the first benchmarks for Jalapeño, the inference chip it built with Broadcom, claiming up to 3.6x lower latency and 1.9x more work per watt than Nvidia’s Blackwell systems.
The 700-watt chip beat Nvidia’s GB200 and GB300, which draw 1,200 and 1,400 watts. OpenAI ran the tests itself on SemiAnalysis’s public InferenceX benchmark, and no outside lab has verified the numbers.
OpenAI won’t sell Jalapeño, keeps buying Nvidia for training, and starts deploying late this year. It also used Codex to help design the chip, so its models are quietly building their own hardware.
Perplexity and Nvidia launch a local AI agent
Perplexity and Nvidia launched Portable Computer, a version of Perplexity’s Computer agent that runs entirely on a user’s own machine. Files stay on the device and local work burns no credits.
It runs Qwen 3.8 27B or Perplexity’s post-trained PPLX 27B, and the agent asks permission before routing any step to one of 15-plus cloud models. Pro and Max subscribers get it free.
The Information reported days ago that Nvidia may invest billions in Perplexity at a $30B-plus valuation. Local-first still starts with buying Nvidia’s $4,699 DGX Spark.
Meta settles child harms lawsuit for $18 billion
Meta agreed to pay up to $18 billion to settle claims from 29 states that Facebook and Instagram harmed children through addictive designs and illegally collected minors’ data.
The deal adds default time limits, overnight blocks, hidden like counts, stricter filters, and stronger age checks for teens, with the protections set to remain in place for 10 years.
Meta tied about $5.3 billion of the settlement to YouTube and TikTok adopting similar safeguards, effectively turning part of its legal bill into pressure for industry-wide rules.
Robot brains still need their ChatGPT moment
Robotics startups have raised billions to build smarter machines, but developers still struggle with limited training data, unreliable manipulation, and models that cannot consistently perform useful real-world tasks.
Builders increasingly rely on simulation, reinforcement learning, and larger datasets, while autonomous vehicle companies like Wayve and Uber expand into humanoid robotics using infrastructure developed for self-driving systems.
Industry leaders compare robotics today to AI’s GPT-2 era, but specialized machines already work in construction and warehouses, suggesting useful robots may arrive vertically before general-purpose humanoids do.