Version 0.1 · effective 13 Sep 2026
Methodology
NetGoodIndex scores discrete real-world events, not marketing claims, benchmark exams, or speculative future risk.
What is scored
An event is a documented beneficial or harmful outcome with a plausible link to one or more AI systems. A press release that says a model “accelerates science” is not an event. A specific result that later shows up in a paper, clinic, court, operational system, or widely deployed product can be.
Significant product or feature launches are in scope when they change professional practice in an industry—GitHub Copilot for software engineering, Claude Code for agentic coding, a chip-layout system used in production TPUs. Download counts, token volume, and “X million users” alone are not scores. The launch still needs evidence of adoption or a measured productivity effect.
Benefit and harm stay on separate ledgers. Net Good is benefit minus harm. A model with large benefit and large harm should not look like a quiet success.
Score formula
Points = Base Impact × Attribution × Evidence × Realization × Durability × Credit Share
Every multiplier is between 0 and 1. Credit shares across AI contributors cannot exceed 1.0.
- Tier I Useful / Limited harm: 1
- Tier II Notable / Significant: 3
- Tier III Major: 10
- Tier IV Historic: 30
- Tier V Civilizational / Catastrophic: 100
Credit units
Organizations, families, and models are three different primitives. The scoring unit is a named model release—GPT-6 Astra, GPT-5.6 Sol, Gemini Deep Think—not a lab and not a product brand.
- Organization is the lab or company (OpenAI, Google DeepMind). Org scores are rollups.
- Family is a named generation or product line (GPT-6, GPT-5.6, Claude, Gemini). Family scores sum the releases inside it.
- Named release is the default model record. Astra is not credited for work a lab says was done by a more capable unreleased system.
- Internal / unreleased is its own model when a lab documents an unnamed or unreleased system. Date it. Do not dump those points onto the public flagship.
- If authors used “Claude” or “Codex” without a version, record a version-unspecified model in that family until a named release is identified.
Credit shares across AI contributors still cannot exceed 1.0. A Lean formalization by Astra of a proof found by an internal system is a split, not a transfer of the whole event to Astra.
Rumors that Hodge, Birch–Swinnerton-Dyer, Riemann, Yang–Mills, or P vs NP have been solved are not events. Poincaré remains the only Clay Millennium problem with an accepted human solution. A lab announcement is recorded as a claim, at most Major until independent reconstruction.
Evidence and review
Developer blogs score as announcements, not proof. Peer review, independent replication, court records, and measured outcomes raise evidence. High-tier events require human review. Every published score change writes an immutable revision.
Retracted events remain on the record with a score of zero unless residual impact is documented. Speculative catastrophe, chatbot quality, token volume, and popularity are out of scope.
Corrections
Anyone can submit a factual correction, missing source, attribution dispute, or retraction notice. Submissions enter an editorial queue. Leaderboard numbers are never typed by hand.