Topic
Multimodal
Vision-language models, audio AI, and cross-modal capabilities
Featured
Google Photos adds AI video editing with cinematic relighting

Google's Gemma 4 12B Brings Multimodal AI to Offline Laptops
All Stories
NVIDIA Opens Alpamayo 2 Super for Commercial AV Use
NVIDIA has released Alpamayo 2 Super, an open-source reasoning model for autonomous vehicles, under a permissive…
Brain Waves Join Video as Physical AI Training Data
Frontier physical AI models are moving beyond video training data to incorporate multiple camera angles, dense…

Black Forest Labs Launches FLUX 3 Video Model in Limited Release
Black Forest Labs launched FLUX 3, a multimodal AI model capable of generating images and video with audio up to 20…
DeepMind Researcher Raises $300M Pre-Seed on Visual AI Vision
Andrew Dai, a former DeepMind researcher who contributed to foundational AI research that influenced ChatGPT's…
Meta Launches Muse Image AI Model Across Social Platforms
Meta has launched Muse Image, an AI image generation model developed by its Superintelligence Labs division, now…
Google Brings Personalized Image Generation to Free Gemini Users
Google is making personalized AI image generation available to eligible free Gemini users in the U.S. The feature…

Multimodal AI turns aerial imagery into searchable data
AWS and Vexcel, an aerial imagery provider operating across 45+ countries, developed a multimodal AI system that…
Google DeepMind's Gemma 4 Now Available on AWS Bedrock
Google DeepMind's Gemma 4 model family is now available on Amazon Bedrock, offering three instruction-tuned variants…
PixelRAG bypasses text parsing, cuts RAG costs 10x
Researchers from UC Berkeley, Princeton, EPFL, and Databricks introduced PixelRAG, a retrieval system that bypasses…
Google DeepMind Releases Gemma 4 12B for Laptop-Based AI
Google DeepMind introduced Gemma 4 12B, a multimodal AI model designed to run on consumer laptops with 16GB of RAM. The…

Google Launches Near Real-Time Voice Translation in Gemini 3.5
Google has launched Gemini 3.5 Live Translate, a near real-time speech translation feature now available in Google AI…

Qwen3.7-Plus Now Available Only Through Alibaba's Proprietary API
Alibaba released Qwen3.7-Plus, a multimodal AI model supporting text, video, and image inputs at $0.40/$1.60 per 1M…

AWS Adds Multimodal Evaluators to Strands Evals
AWS has announced four multimodal evaluators for Strands Evals that use large language models as judges to assess…

Google Merges Search and AI into One Interface
Google has redesigned its search box for the first time in 25 years, transforming it from a simple keyword input field…

Databricks Integrates GPT-5.5 for Enterprise Agents
Databricks has integrated OpenAI's GPT-5.5 model into its enterprise agent workflows following the model's performance…

ML Identifies Cancer Immunotherapy Targets with Patient Validation
Researchers at Augustine et al. have developed a multimodal graph neural network designed to identify cancer…

Pulse AI and Bedrock Cut Financial Document Processing From Days to Hours
AWS and Pulse AI have demonstrated a financial document processing pipeline that combines Pulse's document…

Thinking Machines Previews Full-Duplex AI for Real-Time Conversation
Thinking Machines, the AI startup founded by former OpenAI CTO Mira Murati and researcher John Schulman, has unveiled a…

Amazon Nova Multimodal Embeddings Unlock Cross-Modal Search for Manufacturing Docs
Amazon has released Nova Multimodal Embeddings, a model that maps text, images, and document pages into a shared vector…

OpenAI Adds Reasoning to Realtime Voice Models
OpenAI has released new realtime voice models available through its API that can reason, translate, and transcribe…

Pet Camera Startup Cuts Inference Costs with AWS Inferentia2
Tomofun, maker of the Furbo pet camera, migrated its vision-language model inference from GPU-based EC2 instances to…

ChatGPT Images 2.0 gains traction in India, lags elsewhere
ChatGPT Images 2.0 has gained significant traction among Indian users who are leveraging the tool to create personal…

Google TV Adds Gemini Photo and Video Tools
Google TV is integrating additional Gemini AI features, including photo and video transformation capabilities powered…

Poly-DPO and ViPO: Scaling Visual Preference Optimization
Researchers introduced Poly-DPO, an algorithmic extension to preference optimization that adds a polynomial term to…

NVIDIA Nemotron 3 Nano Omni Consolidates Multimodal AI for Agents
NVIDIA and AWS announced day-zero availability of Nemotron 3 Nano Omni on Amazon SageMaker JumpStart, a…