Claude Filled In the Sky’s Missing Third

Astronomers have never had a complete map of the sky in ultraviolet light. Now they do, and about a third of it was predicted by AI rather than photographed. According to Anthropic, Brice Ménard used Claude Science to build the first full-sky UV map. Ménard is an astrophysicist at Johns Hopkins and also a researcher at Anthropic. The map covers far-UV light at 154 nm and near-UV light at 232 nm.

The interesting part isn’t only the map. It’s how much of the work an AI agent handled, and how carefully the team separated what was measured from what was guessed.

🔭 Why the Map Had a Hole

NASA’s GALEX mission did most of the heavy lifting. It imaged about two-thirds of the sky in UV. It skipped areas with very bright stars on purpose, because they could damage its detectors. That included much of the Milky Way’s plane, which happens to be one of the most scientifically interesting regions.

For years that left a gap in one of astronomy’s basic reference maps. Other instruments captured pieces of the missing sky, but nobody had stitched it all together into one consistent product.

🛠️ How Claude Did It

Anthropic reports that Claude coordinated a team of AI agents to do work that usually eats months of a researcher’s time:

  • Gathering data: The agents downloaded data from six sources: GALEX, Swift, FIMS/SPEAR, TD-1, Planck and Gaia.
  • Calibrating and merging: Each instrument has its own quirks, resolution and units. The agents calibrated the datasets and merged them onto a common footing.
  • Inpainting the gaps: For regions with no UV data, the system used inpainting. That’s a technique that fills in missing parts of an image based on what’s around them. Here, the model learned how UV brightness tracks visible, infrared and radio light, then predicted UV values where only those other wavelengths existed.
  • Refining: The team ran several rounds of improvement to tighten the predictions.

📊 The Numbers

  • Measured share: about two-thirds of the sky, mostly from GALEX
  • Predicted share: about one-third, including much of the galactic plane
  • Accuracy on hidden test data: within about 10% of real UV measurements after refinement
  • Labeling: every pixel is tagged “measured” or “predicted,” with uncertainty estimates

That 10% figure comes from a standard check. The team hid real UV data from the model, asked it to predict those areas, then compared. For an all-sky reference map, that’s a solid start. It isn’t a replacement for actually observing those regions, though.

⚠️ Where Humans Still Mattered

This is the part practitioners should pay attention to. According to reports on the project, the map had an image defect that made it through two rounds of AI review. A human researcher spotted the artifact that the agents missed.

That fits a pattern we’re seeing across agentic science. Agents are great at the slow work of pulling data, cleaning it and running pipelines. They’re less reliable at noticing when something just looks wrong. Domain experts still catch the errors that matter.

A few other caveats:

  • Most of what’s known so far comes from Anthropic’s own blog post, not a peer-reviewed paper.
  • It’s not yet clear whether the full datasets and methods will be published for independent checking.
  • The predicted third is a model estimate. Anyone doing precise work in the galactic plane should treat those pixels with care.

💡 Why This Matters

What stands out here is the design choice to label every pixel by where it came from. Mixing AI-generated data with real measurements is risky in science. Once predicted values slip into a dataset unlabeled, they can quietly contaminate later research. Tagging each pixel and attaching an uncertainty estimate is the responsible way to do it, and other teams should copy it.

If you’re building AI pipelines in any data-heavy field, here’s what to take from this:

  • Use agents for data wrangling. Downloading, calibrating and merging messy sources is where they save the most time.
  • Track where every value came from. Keep a clear line between observed and inferred values.
  • Test against held-out data. Hide real data, predict it, then measure the error.
  • Keep an expert in the review loop. AI review alone missed a visible defect here.

🚀 What Comes Next

The project shows that tools like Claude Science can now take on real research infrastructure, not just literature summaries or code snippets. The next test is outside scrutiny. If the data and methods get released and hold up, expect other teams to try the same approach on other incomplete sky surveys and patchy scientific datasets.

You can find more details in Anthropic’s original post.

Scroll to Top