WorkThe story, in brief

What Do Your Logits Know? (The Answer May Surprise You!)

Apple researchers just proved your 'private' model internals leak everything. Here's what attackers can extract.

Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
People, judgement and the changing nature of work.AI illustration by KeyNews
The KeyNews take

Why it matters

Apple's research reveals that probing model internals can extract sensitive information thought to be inaccessible to users, raising critical security and privacy concerns for deployed vision-language models and exposing a gap in model owner threat models.

The key facts

10 to know
  1. Study focuses on information leakage through model internals (logits, residual streams)

  2. First systematic comparison of information retention across representational levels

  3. Vision-language models used as primary testbed

  4. Identifies both unintentional and malicious information leakage vectors

  5. Published by Apple Machine Learning Research

  6. Challenges assumption that low-dimensional projections adequately protect model information

  7. Research demonstrates systematic information leakage through logits and residual streams

  8. Vision-language models used as testbed for probing model internals

  9. Information retention across representational levels identified as security/privacy risk

  10. Unintentional or malicious information access possible despite model owner assumptions

Go to the source

Apple Machine Learningmachinelearning.apple.com

Publisher excerpt: Recent work has shown that probing model internals can reveal a wealth of information not apparent from the model generations. This poses the risk of unintentional or malicious information leakage, where model users are able to learn information that the model owner assumed was inaccessible. Using…
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work