Topic

Data Quality

5 research articles on Data Quality — drawn from 53 integrated data sources with computational validation.

Kill Report

63 GB of Data That Never Made It to Production

A sync-drift audit revealed 1,117 files (~63 GB) existing only locally, never uploaded to production. 32 common database tables had differing row counts with local exceeding production by 12.6M rows. A DigitalOcean key-naming bug duplicated source prefixes in every object path. Zero data corruption — but massive invisible drift.

Methodology

13 Ways Computational Materials Science Goes Wrong

A catalog of interpretation errors and compute failure modes from our materials discovery pipeline: thermostat overshoot (625K target / 750K actual), MACE artifacts on perovskites, PBE bandgaps presented without HSE correction, metallic DFPT results, phonon instabilities, 53 files with THz vs cm⁻¹ confusion, and more.

Kill Report

The Data Plane Kills Ideas That Market Analysis Can't

Four invention candidates rated 'plausible' by web-search-based competitive analysis were instantly killed when tested against real warehouse data. The patent-clear composition finder's core anti-join produced false results because the canonical crosswalk was never built, patent formulas were stored as un-reduced supercells, and 12.4% of formula values were filesystem paths.

Methodology

NULL Should Never Mean Favorable: How Default-Optimistic Screening Corrupts Results

A data-stitching audit revealed that missing phonon data defaulted to 'stable,' missing patent counts defaulted to 'novel,' and missing PFAS flags defaulted to 'clean' across our screening pipeline. Every screen that treats NULL as favorable silently promotes unvalidated candidates. Here's what we found and how we fixed it.

Methodology

Where DFT Codes Disagree: A Cross-Database Energy Map

Cross-source DFT disagreement analysis across 0 compositions. 0 show high disagreement where independent DFT databases predict substantially different thermodynamic stability.