ToolsThe story, in brief

Gemini API File Search is now multimodal

Gemini API just went multimodal. File search now handles images, PDFs, and video—changing how developers build RAG.

Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
New tools for building and creating with AI.AI illustration by KeyNews
The KeyNews take

Why it matters

Google expands Gemini API capabilities with multimodal file search, enabling developers to build retrieval-augmented generation systems that work across text, images, video, and documents. This lowers the barrier to building sophisticated AI applications and directly competes with similar RAG offerings from competitors.

The key facts

6 to know
  1. Gemini API File Search now supports multimodal inputs

  2. RAG (retrieval-augmented generation) functionality expanded

  3. Support for images, PDFs, and video documents

  4. Published May 10, 2026

  5. Developer-facing feature rollout

  6. Blog post from Google AI official channels

Go to the source

Hacker Newsblog.google

Publisher excerpt: Article URL: Comments URL: Points: 3 # Comments: 0
Read original report
Back to today's editionMore tools news

The wider picture

View all
Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
AI illustration by KeyNews
Tools01

OpenAI nabs key Patreon execs ahead of upcoming announcement

OpenAI is building a creator-focused product suite with deep domain expertise (Patreon's co-founder + product + engineering leads). This signals a major new revenue and engagement vector for ChatGPT — and a direct threat to Patreon's existing creator economy.

The Verge AI
Illustration of a transparent lens revealing connected networks across layers of paper.
AI illustration by KeyNews
Tools02

Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model

A practical, deployment-ready layer for a common production pattern — classification over generation — that practitioners can drop into existing LLM stacks immediately. No fine-tuning required.

MarkTechPost
Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
AI illustration by KeyNews
Tools03

SpeakON Ships a MagSafe AI Voice Button With Its Own Microphone

A hardware-first approach to voice AI workflow — MagSafe button with onboard mic addresses the real friction in voice-to-text-to-action. Relevant to practitioners building voice UX and to the broader consumer AI tooling wave.

MarkTechPost