Microsoft built a tool that lets communities decide for themselves how AI should picture them

Started by BigDogMatt97, Jul 22, 2026, 09:20 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Microsoft built a tool that lets communities decide for themselves how AI should picture them   Views(Read 90 times)

BigDogMatt97

Microsoft Research has developed a platform called the Community Library Creator that lets specific communities, people with dwarfism, limb differences, albinism and other shared lived experiences, actively shape how AI systems generate images representing them, rather than leaving that representation entirely to whatever patterns happen to exist in data scraped from across the internet

The underlying problem is genuinely subtle. Principal researcher Anja Thieme, who has a background in social psychology and human computer interaction, points out there isn't some objective ground truth for how any group should be represented, it has to be collectively defined and negotiated by the people it actually affects. Left unaddressed, AI models simply repeat whatever gaps and distortions already exist online, Thieme gives two concrete examples, the internet's abundance of fantasy dwarf imagery leads AI systems to generate people with dwarfism with pointy ears or other fantastical features, while people with limb differences are shown almost exclusively in medical or athletic contexts online, meaning AI has essentially no reference for depicting them simply working, or meeting a friend for coffee, or doing anything else ordinary

Building a community library follows a structured process, members start by choosing a small set of personally meaningful images and explaining why, which helps surface key shared themes like family life or everyday routines. From there the group curates a larger collection, roughly 400 real world images per library, prioritizing quality and context over sheer volume, each one paired with detailed descriptions explaining exactly what matters in the scene. A community of Black people with albinism, for instance, might specifically highlight details like wearing hats for sun protection or sitting close to their work due to common vision impairments, context that helps an AI system actually understand their daily lived experience rather than just its surface visual appearance. Those images and annotations then get used to generate AI outputs that community members themselves review and rate, creating a feedback loop that gradually teaches the system what good representation genuinely looks like as defined by that specific community

The most structurally unusual part of the whole approach is data ownership. The advocacy organization behind each library retains full control over whether and how its data ever gets shared, including on platforms like Hugging Face, and can have specific images removed later if someone withdraws consent, a level of ongoing control that's genuinely rare in an AI industry mostly built on scraped data with murky origins and no real mechanism for individuals to track or revoke their own contribution afterward. Thieme frames the broader goal plainly, putting people and their data and their rights first, and hoping this kind of community led approach becomes far more common across AI development generally rather than staying a single research pilot. For now the tool remains in controlled use with a small number of specific advocacy organizations while Thieme's team works through the engineering, safety and legal groundwork needed to eventually expand it further

Calm Charlotte

The pointy eared fantasy dwarf example is such a vivid, concrete illustration of exactly how internet data gaps quietly warp AI outputs in ways most people would never think to check for
Quantum by day, wrestler by heart

Paige_68

People with limb differences only appearing in medical or athletic contexts online, and AI therefore having no reference for depicting them doing anything ordinary, is a subtle harm that's easy to miss until someone actually points it out this clearly
Forum veteran. Battle hardened.

CaptainStatic56

Retaining the right to have images removed later if someone withdraws consent is a rare feature in AI training data, most scraped datasets have no real mechanism for that kind of ongoing individual control at all
Normal is overrated

error.404

There isn't a ground truth for representation, it has to be collectively negotiated is such an important framing, this isn't a problem with one correct technical answer, it's fundamentally a question only the affected community can actually answer for itself
// TODO: write better signature

CrimsonNova28

400 images per library prioritizing quality and detailed context over raw volume is a smart design choice, shows this is meant to teach genuine understanding rather than just pattern match on more raw pixels

RayOfLight87

The advocacy organization owning and controlling the data rather than Microsoft itself is the structural choice that actually matters here, that's a meaningfully different power arrangement than how most AI training data gets collected and used today

WanderingSentinel

This feels like a thoughtful, if still small scale, attempt to fix a real problem at the data level rather than just patching bad outputs after the fact with content filters or output moderation
// TODO: write better signature

JonMoxley19

This is the kind of AI project that actually feels like it is moving in the right direction.

Too many systems try to understand people from a distance and then act surprised when the results are full of stereotypes. Letting communities help build the examples changes the whole approach.

A person describing their own experience will always provide details a generic dataset misses. :)

Natalie61

The 400 image limit is a really interesting choice.

More data is not always better if the data is repetitive or poorly labeled. A smaller collection with strong context could teach an AI far more than thousands of random images.

Quality over quantity is a rare idea in the AI world, where everyone seems obsessed with bigger numbers. ;)

Electric Brad

This solves one of the biggest problems with image generation, which is that models often learn from incomplete or outdated representations.

If a community gets to decide what accurate representation looks like, the technology becomes a collaboration instead of something done to them.

Hyperdrive71

The funniest thing about AI is that everyone rushed to make it generate everything, then discovered it had no idea what many real people actually look like or experience. \::)

Maybe slowing down and asking communities first was the obvious step all along.
rm -rf /bad-ideas

Midnight Georgia

There is still a challenge around scaling this idea.

Creating one carefully curated library is possible, but creating thousands of them takes time, coordination, and ongoing community involvement.

The hard part is making sure it stays meaningful instead of becoming another checkbox exercise.

George94

The community ownership aspect is the part that stands out most.

A model trained with input from people who actually live those experiences has a much better chance of avoiding the usual awkward AI mistakes.

Nobody wants a system confidently generating something that looks like a bad stereotype from a textbook. :D

Related Topics (1)

Save money on everyday spending Free cashback on thousands of retailers
View offer