Skip to Content

AI News Update: Can I buy OpenAI’s new Jalapeño chip to run ChatGPT faster on my own hardware? and more

OpenAI’s custom inference chip promises faster ChatGPT and lower costs

OpenAI just showed off real benchmark results for Jalapeño, its first custom chip built specifically to run AI models, not train them. Think of inference as the moment ChatGPT actually writes your reply. That is what this chip is built for.

It was co-built with Broadcom and went from design to finished chip in just nine months, partly because OpenAI used its own AI models to help design it.

Here is what makes it interesting:

  • Targets ~50% lower cost per response compared to Nvidia’s current best chips
  • Handles both speed and volume without trading one for the other, something most chips struggle with
  • Keeps the model’s short-term memory close to the processor, so less data travels around and replies come back faster
  • A Gen 2 is already deep in development, Gen 3 is taking shape

For you, this means faster Codex responses, snappier ChatGPT, and fewer slowdowns as usage grows. You cannot rent or buy this chip. It runs OpenAI’s infrastructure only.

Anthropic gives Claude persistent memory across all chats by default

You know that annoying thing where you explain your whole project setup to Claude in chat, then switch to Claude Cowork (the AI agent that handles multi-step tasks like building docs or running code), and it has zero clue what you just said? That’s fixed.

Claude now shares one unified memory across both chat and Cowork. Context you built up in chat is instantly available when Cowork runs a task, and vice versa. No more rebriefing.

Here’s what you can actually do with it:

  • Say “remember this” mid-chat to save something specific
  • Go to Settings > Memory to read, edit, or delete anything Claude has saved
  • Enable sensitive topics (health, beliefs) manually if you want that context saved too
  • Memory updates live during conversations, not just after they end

It’s on by default on Free, Pro, and Max plans. You’re in full control of what stays.

Perplexity ships fully local AI agent that runs on NVIDIA DGX Spark

Perplexity just shipped Portable Computer, and the core idea is simple: your AI agent runs entirely on your own machine. No cloud. No data leaving your desk.

Until now, every Perplexity agent ran on remote servers. Portable Computer flips that. The orchestrator (the brain that plans tasks), the subagent (the part that executes them), and the full tool harness all run locally on your hardware.

Here is what that unlocks:

  • Zero token costs for local tasks. Long loops, big file analysis, repo-scale work, all free.
  • Your sensitive files never leave your device. PII is flagged automatically.
  • Cloud escalation is opt-in, per step, with your approval each time.
  • Connects to Google Drive, Gmail, GitHub, and Slack.

It runs on the NVIDIA DGX Spark with either PPLX 27B or Qwen 3.8 27B. Setup is one click inside the Perplexity app. Available now for Pro and Max subscribers on Linux.

Anthropic’s $30 Trillion Joke

Anthropic plans to tell investors its total addressable market is $30 trillion, roughly the entire US GDP, according to the Wall Street Journal. The company is preparing for an IPO that could raise up to $100 billion at a valuation around $2 trillion. A $2 trillion price tag needs a $30 trillion TAM to look reasonable. For context, SpaceX’s IPO earlier this year claimed a $28.5 trillion TAM. Anthropic just raised the bid.

SpaceX’s $28.5 trillion TAM was already absurd, but investors could at least point to near-monopolies in launch and satellite internet underneath the fantasy. Anthropic has no equivalent moat. It is asking investors to price a $2 trillion IPO on the assumption that AI will subsume the entire economy. Meanwhile, its revenue growth is decelerating: 58% monthly in April, 38% in July. The company needs investors to ignore the slowdown and focus on the horizon. The horizon is the entire US economy. That is not a forecast.

Anthropic is asking investors to believe a $30 trillion TAM while its real revenue slows by the month. That is not a pitch. It is a number so big it stopped being impressive and started being ridiculous.

MIT Taught AI to Predict the Unseen

MIT engineers published a paper in Nature Communications on August 20 detailing η-learning, an algorithm that generates plausible extreme events without ever being trained on extreme event data. It learns from ordinary daily records, then produces complete spatial maps of unprecedented scenarios: where a 300mm storm would hit, how large, how intense, how long. The method applies to floods, wildfires, financial crashes, and supply chain disruptions.

Most AI learns from the past to predict the future. The problem is that the future keeps producing events the past never recorded. Hurricane Katrina was a 30-year event. What does a 100-year Katrina look like? No dataset can answer that question. η-learning does not need one. It generates realistic worst-case maps from the statistical structure of normal data. The jump is not about weather. It is about what happens when every industry that prices risk, insurance, infrastructure, supply chains, financial markets, can suddenly see the disaster before it exists. That is not a better model. That is a different kind of intelligence.

AI spent its entire history learning the past. η-learning just learned the future. Every industry that depends on knowing the worst-case scenario may soon get a tool that does not need a past to draw the picture.

CUDA “Died” Again

OpenAI just unveiled its first custom chip, Jalapeño, a 700W inference accelerator that outperforms Nvidia’s GB200 and GB300 on tasks up to 104x faster. Built with Broadcom, designed in nine months, and benchmarked by SemiAnalysis, Jalapeño costs $1.56 per chip-hour and is the first of three planned generations. SemiAnalysis declared that “CUDA’s moat may be dead.”

CUDA has been declared dead before. Google TPU was supposed to kill it. Google kept TPU for itself. Groq was supposed to kill it. Nvidia bought Groq’s technique for $20 billion. Now OpenAI’s Jalapeño is the latest executioner. It is fast, it is efficient, and it will probably never be sold to anyone outside OpenAI. Every chip that “kills CUDA” ends up either acquired by Nvidia or locked inside a single company’s data center. The grave gets dug, the ceremony is held, and the mourners go back to buying Nvidia GPUs.

CUDA has been declared dead multiple times. It is still alive, and the companies that declared it dead are still sending Nvidia checks. The only thing Jalapeño killed may be the idea that anyone outside OpenAI will ever get to use it.

Phlebotomists Just Got Replaced

Vitestro, a Dutch company founded in 2017, received FDA De Novo authorization this month for Aletta, the first robotic device cleared to draw blood autonomously. Near-infrared light maps the arm, ultrasound scans for veins, and a computer vision model trained on annotated ultrasound images picks the target and guides the needle. In clinical trials, Aletta hit a 95% first-stick success rate, a 0.6% haemolysis rate, and a median draw time of 1 minute 49 seconds; and one phlebotomist can supervise three devices.

Every hospital draws blood. Nobody wants the job. The US has 140,000 phlebotomists, 18,000 openings a year, and people quit faster than they can be replaced. Aletta does not quit. It does not miss more veins on darker skin. It does not get tired. It took eight years of clinical trials, and the FDA wrote a rulebook so competitors can follow. The most boring robot in the world just solved the most boring problem in medicine, and every hospital administrator who has ever spent a morning trying to fill a phlebotomist shift knows exactly what that is worth.

The blood draw is the most common procedure in medicine, and the people who do it are leaving faster than they can be replaced. Aletta does one thing. It does it better than a human. And it never quits.

Spacex Plans $100B Louisiana Spaceport

SpaceX unveiled plans to build the world’s largest spaceport in Louisiana with 10 launchpads, boasting a combined launch capacity of 15,000+ launches per year — more than twice all rocket launches in history. The $100 billion facility would make Louisiana the new epicenter of commercial spaceflight.

Perplexity Launches Portable Computer AI Agent

Perplexity AI introduced a portable computer running an on-device AI agent, pushing the boundary of where AI inference happens. The device brings Perplexity’s search and reasoning capabilities to standalone hardware.

Waymo to Launch Robotaxis in Germany by 2027

Waymo announced plans to launch driverless rides in Germany in 2027, marking its third international market outside the U.S. The Alphabet-owned autonomous driving company continues its measured global expansion strategy.

Anthropic Updates Claude Memory for Customization

Anthropic updated Claude’s memory capabilities, allowing the AI assistant to better retain user preferences and context across sessions while adding protections for sensitive topics. The update improves personalization without compromising safety guardrails.

OpenAI Launches Admin Plugin for ChatGPT Work

OpenAI announced the Admin plugin for ChatGPT Work and Codex, giving enterprise administrators centralized control over AI tool deployments, usage policies, and compliance monitoring — a critical step for enterprise AI adoption.

Google Launches Gemini for Legal Work

Google launched Gemini for legal work, a specialized version of its AI model designed to automate contract analysis, legal research, and document review. The move targets the $40 billion legal tech market, competing directly with specialized AI legal startups.

Apple Announces Mac Mini with M6 and M5 Pro

Apple unveiled the next-generation Mac Mini featuring M6 and M5 Pro chips, continuing its silicon transition with significant performance gains. The compact desktop targets both consumers and professionals.

Meta’s Paid AI Agent Hatch Launches Soon

Meta’s paid AI agent Hatch is launching soon alongside a new model codenamed Watermelon due in October. The move marks Meta’s push into monetized AI agents, competing with OpenAI’s premium tiers and Anthropic’s enterprise offerings.

Nvidia Doubles Edge Robotics Compute with Jetson

Nvidia doubled the compute performance of its entry-level edge robotics platform with the Jetson Orin Nano 2, enabling more sophisticated AI inference directly on robots without cloud connectivity. The upgrade targets the booming autonomous machine market.

ChatGPT Work Can Now Sign Into Websites Automatically

ChatGPT Work gained the ability to use its computer and browser to securely sign into websites on behalf of users without ever seeing their credentials. The feature unlocks dozens of automation use cases from booking appointments to filing insurance claims.

OpenAI Finishes ‘Bel’ Pretrain — GPT-6 Base

OpenAI completed its next massive pretraining run codenamed ‘Bel,’ a successor to ‘Doug’ with over 10 trillion total parameters — comparable to GPT-4.5 in scale. Bel is expected to serve as the base for the Astra assistant and GPT-6 after further RL.