This Week in AI · Jun 20–Jun 26, 2026
A wave of new computer-use features and specialized hardware is reshaping how builders integrate AI into everyday workflows.
What shifted
Gemini 3.5 Flash Adds Native Computer Use
Google · 24 June
Google's Gemini 3.5 Flash now includes built-in computer use, letting the model interact with browsers, mobile apps, and desktop interfaces directly from prompts. For AI builders, this means a single LLM endpoint can stand in for a stack of specialized automation tools, simplifying integration for small businesses and content creators who don't want to wire together five different APIs.
OpenAI & Broadcom Unveil LLM-Optimized Jalapeno Chip
OpenAI · 24 June
OpenAI and Broadcom have jointly released the Jalapeno chip, a custom silicon design optimized for large-language-model inference. The board is intended for internal use at OpenAI but signals a broader industry trend toward dedicated AI hardware that can lower latency and cost compared to general-purpose GPUs.
NVIDIA & AWS Bring GPU Acceleration to Production Scale
NVIDIA
NVIDIA has partnered with Amazon Web Services to embed its GPU acceleration into AWS OpenSearch and EC2 services. The collaboration delivers higher price-performance, lower latency inference, and built-in vector search capabilities directly within the cloud platform. This reduces the operational burden for enterprises deploying large language or vision models at scale.
Apple Skips M6 in Favor of AI-Capable M7 Series
Bloomberg · 25 June
Apple will launch the new M7 Pro, Max, and Ultra chips, with meaningful AI capability improvements, instead of its high-end M6 line. The move embeds large-model inference capabilities into mainstream Macs, potentially reducing reliance on cloud services for tasks such as image editing, transcription, and real-time translation.
Also this week
- openai: How agents are transforming work — Agents directly affect how everyday users can automate multi-step tasks, offering tangible workflow improvements that Ultra Prompt can help them adopt.
- deepmind: Introducing computer use in Gemini 3.5 Flash — The release of a faster and cheaper Gemini model directly impacts user workflows and costs, making it a critical topic for Ultra Prompt to address with practical advice.
- platformer: The CEO of AWS on why Amazon is hiring 11,000 interns and junior employees — The shift directly affects everyday AI users by offering new productivity tools that can be adopted immediately in their day-to-day operations.
- hn: Mistral OCR 4 — The release directly impacts everyday AI users by offering a cost-effective, privacy-preserving OCR solution that can be quickly adopted in real workflows.
- stratechery: Memory Chips and China, Microsoft and Chinese Models — The hardware and model sourcing shift directly impacts everyday AI users' cost and performance, making it a high-value story for Ultra Prompt readers.
What it means
The week underscored a trend toward tighter integration of tool use within LLMs and the rise of specialized silicon to support that capability. Builders now have an expanded set of options: Gemini 3.5 Flash's computer-use capability can replace multiple automation APIs, Jalapeno may lower inference costs for OpenAI users, AWS's GPU-powered services simplify deployment at scale, and Apple's M7 line brings meaningful local inference power to mainstream Macs. Amazon's decision to bring on 11,000 interns and junior employees is a signal that the demand for people who can actually work with these tools is outpacing the supply. Attention should turn to how these shifts affect latency, cost per token, and the security of data when moving from cloud to on-prem or hybrid models.