From 16c740a98f58271fc7bb4c6df253b3813c4c1a3e Mon Sep 17 00:00:00 2001 From: nicweyand Date: Sun, 20 Sep 2026 11:57:47 -0400 Subject: [PATCH 1/5] Publish first signed Site Registry catalog trust --- README.md | 10 +-- docs/CONSUMERS.md | 19 +++--- docs/INDEX.md | 1 + docs/PUBLIC_CATALOG.md | 62 +++++++++++++++++++ trust/public-catalog-20260920/POLICY.md | 15 +++++ trust/public-catalog-20260920/policy.json | 23 +++++++ .../publisher-allowed-signers | 1 + .../reviewer-allowed-signers | 1 + 8 files changed, 120 insertions(+), 12 deletions(-) create mode 100644 docs/PUBLIC_CATALOG.md create mode 100644 trust/public-catalog-20260920/POLICY.md create mode 100644 trust/public-catalog-20260920/policy.json create mode 100644 trust/public-catalog-20260920/publisher-allowed-signers create mode 100644 trust/public-catalog-20260920/reviewer-allowed-signers diff --git a/README.md b/README.md index 37917aa..474e009 100644 --- a/README.md +++ b/README.md @@ -12,10 +12,12 @@ facebook -> Facebook (Wikidata Q355) -> https://www.facebook.com/ public suffix: com ``` -The repository contains the library, CLI, schemas, migrations, and synthetic -fixtures. It does not contain a preapproved production dataset. A publisher must -import source evidence, collect signed reviews, and distribute a signed registry -generation. +The repository contains the library, CLI, schemas, migrations, synthetic +fixtures, and independently authenticated public trust roots. It does not place a +mutable production database in Git. Publishers import source evidence, collect +signed reviews, and distribute immutable signed registry generations. The first +public catalog generation is available as a release asset; see +[Public catalog](docs/PUBLIC_CATALOG.md). ## The basic idea diff --git a/docs/CONSUMERS.md b/docs/CONSUMERS.md index ad0c830..20ddad9 100644 --- a/docs/CONSUMERS.md +++ b/docs/CONSUMERS.md @@ -101,12 +101,15 @@ revocation history. Never mutate a complete generation to migrate it. See ## Argand integration `UPSTREAM.json` records the original Argand extraction baseline and file hashes. -The last recorded downstream integration replaced Argand's embedded crate with -signed v0.3.0 revision `ac8282093d8a815c6227cff86e1f40714d510bcd` at Argand -commit `d9dfd1585ce21d9c4136bcc24fa01fe3bfb8ed6e`. +Argand pins signed v0.5.0 revision +`3d3e08cdfd303df9fbd347a9bab2ba52ad575759`. The public beta uses Site Registry as +Navigate's authoritative auto-route catalog. Its native `navigation-catalog/v2` +file is only a collection- and content-policy-bound serving projection compiled +from one exact registry generation; it is not a second independently curated +destination catalog. -Version 0.5 is handed off as a signed standalone revision. Argand should update its -full Git `rev` in a separate coordinated source/build window, compare contract -changes, and rerun navigation compiler, native resolver, API, abstention, -revocation and clean-process gates. Changing the code dependency does not activate -a registry generation or approve a public destination. +Argand source commit `564ee5fc2fa0974a7b0557a914f274bbd4ab654c` records that boundary and the first +public-beta activation. Changing the code dependency alone still does not activate +a data generation or approve a destination. Every downstream must verify the +signed generation, preserve abstentions, apply its own safety policy, and bind any +serving projection to its own eligible corpus or directory policy. diff --git a/docs/INDEX.md b/docs/INDEX.md index 747ac35..d2f7fb7 100644 --- a/docs/INDEX.md +++ b/docs/INDEX.md @@ -8,6 +8,7 @@ - [Migrating to 0.4](MIGRATING-0.4.md): writer migration and trust transition. - [Source licenses](../crates/argand-site-registry/LICENSE_SOURCES.md): exact terms and attribution. - [Consumers](CONSUMERS.md): Rust, Python/CLI, data distribution and Argand transition. +- [Public catalog](PUBLIC_CATALOG.md): download, independent trust roots, verification, scope and refresh contract. - [Trust](TRUST.md): enforced checks and publisher/consumer responsibilities. - [Publishing](PUBLISHING.md): reviewer keys, candidate acceptance and activation. - [Evaluation](EVALUATION.md): bounded JSONL judgments and result interpretation. diff --git a/docs/PUBLIC_CATALOG.md b/docs/PUBLIC_CATALOG.md new file mode 100644 index 0000000..dbbdf89 --- /dev/null +++ b/docs/PUBLIC_CATALOG.md @@ -0,0 +1,62 @@ +# Public signed catalog + +The v0.5.0 Forgejo release publishes the first immutable data generation that any +Site Registry consumer can verify and resolve: + +- release: +- asset: `argand-site-registry-catalog-20260920-v1.tar.gz` +- asset SHA-256: + `d878fa057397effa5dc729d2fa3a689c8edd1f4112ef1326dd6131b3fdeab63e` +- generation pin: + `ede14746da8817aafdf705dd88cfeabbe8d23e1991e43a304acd8eca9249b18a` + +The release also carries a checksum file and an OpenSSH signature under namespace +`argand-site-registry-release`. Verify it against +[`trust/public-catalog-20260920/publisher-allowed-signers`](../trust/public-catalog-20260920/publisher-allowed-signers). +The signed Git history is the independent channel for the trust root; do not learn +the only trusted key from the archive it authenticates. + +```bash +sha256sum --check argand-site-registry-catalog-20260920-v1.tar.gz.sha256 +ssh-keygen -Y verify \ + -f trust/public-catalog-20260920/publisher-allowed-signers \ + -I argand-site-registry-publisher-v1 \ + -n argand-site-registry-release \ + -s argand-site-registry-catalog-20260920-v1.tar.gz.sig \ + < argand-site-registry-catalog-20260920-v1.tar.gz +``` + +After extraction, verify every member with `SHA256SUMS`, then authenticate the +generation and exact reviewer trust root: + +```bash +argand-site-registry activate \ + --generation public-release-v0.5.0/catalog \ + --current current.json \ + --allowed-signers trust/public-catalog-20260920/publisher-allowed-signers \ + --allowed-reviewers trust/public-catalog-20260920/reviewer-allowed-signers \ + --identity argand-site-registry-publisher-v1 + +argand-site-registry resolve \ + --generation public-release-v0.5.0/catalog \ + --pin ede14746da8817aafdf705dd88cfeabbe8d23e1991e43a304acd8eca9249b18a \ + --query "yahoo mail" +``` + +## Scope and trust + +This first catalog is deliberately small. Its disclosed policy uses one automated +evidence-gate reviewer group rather than claiming human-review quorum. Fresh exact +endpoint observations are required, and source conflicts or dangerous drift need +two groups, so the single automated reviewer must abstain on those risks. Sticky +revocations and publisher/reviewer key separation remain enabled. + +Consumers decide whether this policy is appropriate for their use. Preserve typed +abstentions, retain attribution, and apply independent malware and content policy. +Do not route to the first raw lookup result. High-risk or disputed catalogs should +use the unchanged two-human-reviewer reference policy. + +The generation's approvals expire. Installing an immutable archive is not a promise +that every decision stays valid forever: use the resolver's requested time, +consume cumulative signed revocation feeds when published, and move to a newly +signed full generation before relying on renewed decisions. diff --git a/trust/public-catalog-20260920/POLICY.md b/trust/public-catalog-20260920/POLICY.md new file mode 100644 index 0000000..3617480 --- /dev/null +++ b/trust/public-catalog-20260920/POLICY.md @@ -0,0 +1,15 @@ +# Argand automated high-confidence navigation policy + +This generation is a machine-reviewed public navigation directory. It does not +claim two independent human reviewers. The dedicated reviewer identity approves +only exact, unambiguous name and official-site assertions after source evidence +and a fresh bounded endpoint observation are present. + +The policy keeps sticky revocations, requires separate reviewer and publisher +keys, blocks source conflicts and dangerous drift, and gives those risk classes +a two-group threshold that this automated identity cannot satisfy. Ambiguous, +conflicting, stale, unobserved, expired, or revoked routes therefore abstain. + +Consumers choose whether to trust this publisher and policy. The stricter +two-human-reviewer reference policy remains unchanged and available for +high-risk, disputed, or manually governed catalogues. diff --git a/trust/public-catalog-20260920/policy.json b/trust/public-catalog-20260920/policy.json new file mode 100644 index 0000000..0b25b1a --- /dev/null +++ b/trust/public-catalog-20260920/policy.json @@ -0,0 +1,23 @@ +{ + "schema": "argand.site-policy/v1", + "name": "argand-automated-high-confidence-navigation-v1", + "names": { "approvals": 1, "groups": 1 }, + "edges": { "approvals": 1, "groups": 1 }, + "equivalences": { "approvals": 2, "groups": 2 }, + "sticky_revocations": true, + "publisher_reviewer_separation": true, + "require_name_votes": true, + "allow_legacy_reviews": false, + "maximum_approval_days": 30, + "require_edge_observation": true, + "maximum_observation_age_days": 7, + "block_source_conflicts": true, + "block_dangerous_drift": true, + "reviewer_groups": { + "argand-evidence-gate-v1": "argand-automated-evidence" + }, + "risk_thresholds": { + "source_conflict": { "approvals": 2, "groups": 2 }, + "dangerous_drift": { "approvals": 2, "groups": 2 } + } +} diff --git a/trust/public-catalog-20260920/publisher-allowed-signers b/trust/public-catalog-20260920/publisher-allowed-signers new file mode 100644 index 0000000..3d71ab0 --- /dev/null +++ b/trust/public-catalog-20260920/publisher-allowed-signers @@ -0,0 +1 @@ +argand-site-registry-publisher-v1 ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIGfQ/Nk4eQsi7rwhlS3K9/P6vZ+6IZZka2V62iUfKOlB argand site registry publisher 2026-09-20 diff --git a/trust/public-catalog-20260920/reviewer-allowed-signers b/trust/public-catalog-20260920/reviewer-allowed-signers new file mode 100644 index 0000000..170977e --- /dev/null +++ b/trust/public-catalog-20260920/reviewer-allowed-signers @@ -0,0 +1 @@ +argand-evidence-gate-v1 ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAICT/p/gmy3xn+X9H34+aDxW3ss725jn1Ugr+k9dAPMju argand automated evidence reviewer 2026-09-20 From 3e0cc1bc503564fda84a161cd82eebf7d5c5289f Mon Sep 17 00:00:00 2001 From: nicweyand Date: Tue, 22 Sep 2026 08:25:24 -0400 Subject: [PATCH 2/5] release: add Web Graph authority evidence for v0.6 --- CHANGELOG.md | 19 ++ Cargo.lock | 8 +- Cargo.toml | 2 +- README.md | 5 +- .../argand-site-registry/LICENSE_SOURCES.md | 7 + crates/argand-site-registry/README.md | 30 +++ .../argand-site-registry/examples/update.toml | 13 + .../src/adapters/common_crawl_web_graph.rs | 150 +++++++++++ .../argand-site-registry/src/adapters/mod.rs | 10 +- crates/argand-site-registry/src/cli.rs | 8 + crates/argand-site-registry/src/download.rs | 53 +++- crates/argand-site-registry/src/model.rs | 21 ++ crates/argand-site-registry/src/release.rs | 1 + crates/argand-site-registry/src/store.rs | 70 ++++- crates/argand-site-registry/src/update.rs | 30 ++- .../argand-site-registry/tests/common/mod.rs | 3 + crates/argand-site-registry/tests/webgraph.rs | 239 ++++++++++++++++++ docs/CONSUMERS.md | 7 +- docs/FORMATS.md | 12 +- docs/SOURCE-CANDIDATES.md | 3 +- docs/VALIDATION.md | 22 ++ 21 files changed, 698 insertions(+), 15 deletions(-) create mode 100644 crates/argand-site-registry/src/adapters/common_crawl_web_graph.rs create mode 100644 crates/argand-site-registry/tests/webgraph.rs diff --git a/CHANGELOG.md b/CHANGELOG.md index a8314c9..8d83e1a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,24 @@ # Changelog +## 0.6.0 - 2026-09-22 + +- Add a streaming Common Crawl domain Web Graph adapter for harmonic-centrality, + PageRank, and member-host evidence while preserving exact provider fields and + source-line coordinates. +- Authenticate and validate the complete rank stream but retain only registrable + domains already asserted by imported public identity sources, avoiding a + multi-gigabyte runtime catalog whose unrelated rows cannot resolve routes. +- Bind every compact graph projection to the SHA-256 of its sorted candidate + domains and fail closed when identity evidence, the PSL, the bound scope, or a + selected graph row is absent. +- Teach scheduled updates to import identity sources before automatically binding, + downloading, and importing Web Graph evidence with `{candidate_domains}`. +- Strictly allowlist official HTTPS domain-rank objects and retain Common Crawl + Terms-of-Use attribution without treating authority as ownership, safety, + reviewer approval, or query popularity. +- Update Rustls to 0.23.45, remediating RUSTSEC-2026-0285 in the dataset + acquisition path. + ## 0.5.0 - 2026-09-13 - Add a streaming ROR 2.1 ZIP adapter with exact schema checks, declared domains, diff --git a/Cargo.lock b/Cargo.lock index 0734765..afcb5e2 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -78,14 +78,14 @@ checksum = "330a5ed07fa54e4702c9d6c4174f74427fc0ef6e214bbd677ae50a5099946470" [[package]] name = "argand-atomic" -version = "0.5.0" +version = "0.6.0" dependencies = [ "tempfile", ] [[package]] name = "argand-site-registry" -version = "0.5.0" +version = "0.6.0" dependencies = [ "anyhow", "argand-atomic", @@ -1513,9 +1513,9 @@ dependencies = [ [[package]] name = "rustls" -version = "0.23.43" +version = "0.23.45" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0283386ce02abc0151e1761d08802dfe86c173b0b494af5cbc086574e453da06" +checksum = "0d41d731c7d2f962d1ccc364cec258de3c0e93b38c2fb3ba97ac74513048d634" dependencies = [ "aws-lc-rs", "once_cell", diff --git a/Cargo.toml b/Cargo.toml index 510af40..6bfc5f0 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -4,7 +4,7 @@ resolver = "3" members = ["crates/argand-atomic", "crates/argand-site-registry"] [workspace.package] -version = "0.5.0" +version = "0.6.0" authors = ["Nic Weyand"] edition = "2024" license = "AGPL-3.0-or-later" diff --git a/README.md b/README.md index 474e009..93b18eb 100644 --- a/README.md +++ b/README.md @@ -164,10 +164,13 @@ CLI and preserves its JSON contract. | Chrome UX Report | origin popularity bucket, month, optional audience country | CC BY 4.0 International | | Curlie | site titles, categories, descriptions retained for audit | CC BY 3.0 Unported | | Public Suffix List | ICANN and PRIVATE suffix rules | MPL 2.0 | +| Common Crawl Web Graph | domain harmonic-centrality/PageRank and member-host count | Common Crawl Terms of Use (`LicenseRef-Common-Crawl-Terms-of-Use`) | Popularity never proves identity or ownership. Curlie attribution applies to names and categories as well as descriptions; copied descriptions are redacted -unless the caller explicitly exports them and satisfies the display obligations. +from compact display surfaces unless the caller explicitly exports them and +satisfies the display obligations. The scheduled updater can bind the Web Graph +to the exact public-identity domain frontier; it never approves or activates routes. Read [LICENSE_SOURCES.md](LICENSE_SOURCES.md) before distributing provider data. Cloudflare Radar, default Tranco, Cisco Umbrella, arbitrary mirrors, and sources diff --git a/crates/argand-site-registry/LICENSE_SOURCES.md b/crates/argand-site-registry/LICENSE_SOURCES.md index 6db14d7..849d624 100644 --- a/crates/argand-site-registry/LICENSE_SOURCES.md +++ b/crates/argand-site-registry/LICENSE_SOURCES.md @@ -14,6 +14,7 @@ listing is evidence of an assertion, not a guarantee of ownership or safety. | Chrome UX Report (CrUX), Google | [CC BY 4.0 International](https://creativecommons.org/licenses/by/4.0/), [`CC-BY-4.0`](https://developer.chrome.com/docs/crux/methodology) | [Monthly BigQuery dataset](https://developer.chrome.com/docs/crux/bigquery/): `origin`, `experimental.popularity.rank`, observation month, optional audience-country dataset code. The adapter produces `origin,rank,yyyymm,country_code` CSV. Rank is a coarse bucket, not a precise visit count. Audience country is not website jurisdiction. No API key or OAuth token is retained. | | Curlie | [CC BY 3.0 Unported](https://creativecommons.org/licenses/by/3.0/), [`CC-BY-3.0`](https://curlie.org/docs/en/license.html), including the attribution placement prescribed on that page | [Format documentation](https://curlie.org/docs/en/rdf.html), [official download redirect](https://curlie.org/directory-dl), currently [Passau-hosted archive](https://share.innkube.fim.uni-passau.de/curlie-rdf/curlie-rdf-all.tar.gz). Despite its RDF name, the current archive contains **literal TSV**. Content: URL, title, description, category ID. Structure: category ID, full category path, entry count, description, latitude, longitude. Archive notices are retained. | | Public Suffix List contributors | [Mozilla Public License 2.0](https://mozilla.org/MPL/2.0/), [`MPL-2.0`](https://publicsuffix.org/list/public_suffix_list.dat) | [Official list](https://publicsuffix.org/list/public_suffix_list.dat). All ICANN and PRIVATE rules, wildcard/exception rules, version/commit comments and notices. Used for hostname, registrable-domain and public-suffix derivations. Download at most once per day. | +| Common Crawl Web Graph | [Common Crawl Terms of Use](https://commoncrawl.org/terms-of-use), `LicenseRef-Common-Crawl-Terms-of-Use` | [Official Web Graph releases](https://index.commoncrawl.org/web-graphs-index.html). The domain-rank adapter consumes the exact six-column rank file: harmonic-centrality rank/value, PageRank rank/value, reversed registered domain, and provider `n_hosts`. `n_hosts` is retained as `member_hosts`; it is not represented as inbound-linking hosts. This is authority/popularity evidence only and never establishes entity ownership, query popularity, safety, or route approval. | | Argand candidate observer | [CC0 1.0 Universal](https://creativecommons.org/publicdomain/zero/1.0/), `CC0-1.0` | Local host-side observations authored by the registry publisher: HTTP status and redirect targets, canonical/hreflang/JSON-LD/sitemap/country-selector targets, public DNS-set hash, TLS leaf-certificate hash, bounded failure class, content hash and selectors. These records describe a capture; they do not incorporate page prose or prove ownership. | ## Attribution and distribution @@ -56,6 +57,12 @@ listing is evidence of an assertion, not a guarantee of ownership or safety. and `facts`; exports include the PSL fact and original download locator. Changes to covered PSL source files must remain available under MPL 2.0. The Rust `publicsuffix` parser is MIT/Apache-2.0; that is separate from the list. +* **Common Crawl Web Graph:** retain the exact release and object URL, retrieval + time, content digest, Terms of Use link, and identify Argand's reversed-domain + projection. Common Crawl's Terms are not an SPDX open-data license and may + change; re-review them for every new acquisition. The rank file describes a + crawl-derived graph and does not transfer rights in crawled pages. Do not use + rank alone to assert ownership, safety, trust, or an official destination. * **Argand observer:** locally produced observation metadata is dedicated under CC0. The fetched page remains subject to its own rights. The default observer stores only a bounded body in the local replay cache and emits normalized link, diff --git a/crates/argand-site-registry/README.md b/crates/argand-site-registry/README.md index dc412bd..524d327 100644 --- a/crates/argand-site-registry/README.md +++ b/crates/argand-site-registry/README.md @@ -230,6 +230,36 @@ The month belongs in the source snapshot identity. With scheduled typed supersession enabled, the next month then replaces the same audience partition instead of accumulating stale popularity facts. +### Common Crawl Web Graph + +The adapter accepts the official domain-level rank object whose exact header is: + +```text +#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts +``` + +It reverses `host_rev` (`com.facebook` to `facebook.com`) and retains harmonic +centrality, PageRank and `n_hosts` as source-separated popularity evidence. +`n_hosts` means hosts belonging to the registered domain and is exposed as +`member_hosts`; it is not an inbound-link count. The adapter creates no entity, +name, official-site edge, review, vote or route. + +Only exact domain-rank objects below the official +`data.commoncrawl.org/projects/hyperlinkgraph//domain/` hierarchy are +allowlisted. Use `compression: gzip`, the exact release ID as the snapshot, a +positive byte ceiling, and typed coverage. The downloader and importer hash and +consume the complete object; a byte-range prefix must not be declared as the +complete source. Common Crawl's Terms of Use are not an SPDX open-data license, +so preserve the Terms link and re-review it on each acquisition. + +Domain-rank files contain tens of millions of rows. Import the PSL and public +identity sources first, then run `web-graph-selection --database ...` to obtain +the exact `candidate-domains:` scope. A scheduled update may instead use +the `{candidate_domains}` token. The importer still parses and authenticates the +complete stream but persists only matching registrable domains, with original +row coordinates. This keeps the public updater reproducible and compact. Rank +still cannot replace reviewer quorum or destination-safety checks. + ## Build, inspect and review The strict default requires an independently maintained OpenSSH reviewer trust diff --git a/crates/argand-site-registry/examples/update.toml b/crates/argand-site-registry/examples/update.toml index 881b918..abfda29 100644 --- a/crates/argand-site-registry/examples/update.toml +++ b/crates/argand-site-registry/examples/update.toml @@ -56,6 +56,19 @@ collection = "default" kind = "full" supersedes = [] +# Optional Common Crawl Web Graph authority evidence. The updater imports +# identity sources first, replaces this token with their exact sorted-domain +# digest, authenticates the complete graph stream, and retains only matching +# registrable domains. Popularity never authorizes a redirect. +# [[downloads]] +# source = "common_crawl_web_graph" +# format = "common_crawl_domain_ranks_tsv" +# compression = "gzip" +# url = "https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2022-may-jun-aug/domain/cc-main-2022-may-jun-aug-domain-ranks.txt.gz" +# snapshot = "cc-main-2022-may-jun-aug" +# scope = "{candidate_domains}" +# maximum_bytes = 3000000000 + # Optional pinned acquisitions; repeat [[inputs]] for each source. # [[inputs]] # input = "/data/source-object.gz" diff --git a/crates/argand-site-registry/src/adapters/common_crawl_web_graph.rs b/crates/argand-site-registry/src/adapters/common_crawl_web_graph.rs new file mode 100644 index 0000000..71c835c --- /dev/null +++ b/crates/argand-site-registry/src/adapters/common_crawl_web_graph.rs @@ -0,0 +1,150 @@ +// By Nic Weyand! +//! Streaming Common Crawl domain-rank projection; graph rank never creates identity. + +use super::{RecordSink, SourceAdapter, bounded_line}; +use crate::model::{Fact, Record}; +use anyhow::ensure; +use serde_json::json; +use std::{collections::BTreeSet, io::BufRead}; + +const HEADER: [&str; 6] = [ + "#harmonicc_pos", + "#harmonicc_val", + "#pr_pos", + "#pr_val", + "#host_rev", + "#n_hosts", +]; + +pub(super) struct DomainRanks { + pub(super) targets: BTreeSet, +} + +impl SourceAdapter for DomainRanks { + fn ingest(&self, input: &mut dyn BufRead, sink: &mut dyn RecordSink) -> anyhow::Result<()> { + let mut line = String::new(); + ensure!(bounded_line(input, &mut line)? > 0, "empty domain-rank TSV"); + ensure!(columns(&line)? == HEADER, "domain-rank TSV schema changed"); + + ensure!( + !self.targets.is_empty(), + "Web Graph candidate-domain selection is empty" + ); + let mut source_row = 0_u64; + let mut emitted = 0_u64; + while bounded_line(input, &mut line)? > 0 { + ensure!(!line.trim().is_empty(), "blank domain-rank row"); + let fields = columns(&line)?; + ensure!( + fields.len() == HEADER.len(), + "domain-rank column count changed" + ); + let harmonic_rank = positive_integer(fields[0], "harmonic rank")?; + let harmonic_value = nonnegative_finite(fields[1], "harmonic value")?; + let pagerank_rank = positive_integer(fields[2], "PageRank rank")?; + let pagerank_value = nonnegative_finite(fields[3], "PageRank value")?; + let target = reverse_domain(fields[4])?; + let member_hosts = positive_integer(fields[5], "member host count")?; + source_row += 1; + if !self.targets.contains(&target) { + continue; + } + emitted += 1; + let raw = json!({ + "harmonicc_pos": fields[0], + "harmonicc_val": fields[1], + "pr_pos": fields[2], + "pr_val": fields[3], + "host_rev": fields[4], + "n_hosts": fields[5], + }); + sink.emit(Record { + native_id: format!("row:{source_row}"), + raw, + facts: vec![Fact { + subject: target.clone(), + predicate: "popularity".into(), + value: json!({ + "target": target, + "target_kind": "hostname", + "harmonic_rank": harmonic_rank, + "harmonic_value": harmonic_value, + "pagerank_rank": pagerank_rank, + "pagerank_value": pagerank_value, + // Provider n_hosts counts hosts belonging to this domain. It is + // not a count of distinct domains or hosts linking to the target. + "member_hosts": member_hosts, + "country_code": null, + "period": null, + }), + selector: format!("row:{source_row}"), + confidence: 10_000, + }], + })?; + } + ensure!(source_row > 0, "empty domain-rank dataset"); + ensure!( + emitted > 0, + "Web Graph contains none of the selected candidate domains" + ); + Ok(()) + } +} + +fn columns(line: &str) -> anyhow::Result> { + let line = line.strip_suffix('\n').unwrap_or(line); + let line = line.strip_suffix('\r').unwrap_or(line); + ensure!( + !line + .chars() + .any(|character| character.is_control() && character != '\t'), + "control in domain-rank row" + ); + Ok(line.split('\t').collect()) +} + +fn positive_integer(value: &str, field: &str) -> anyhow::Result { + let parsed: u64 = value.parse()?; + ensure!(parsed > 0, "{field} must be positive"); + Ok(parsed) +} + +fn nonnegative_finite(value: &str, field: &str) -> anyhow::Result { + let parsed: f64 = value.parse()?; + ensure!( + parsed.is_finite() && parsed >= 0.0, + "{field} must be finite and nonnegative" + ); + Ok(parsed) +} + +fn reverse_domain(value: &str) -> anyhow::Result { + ensure!( + !value.is_empty() && value.len() <= 253 && value == value.trim(), + "invalid reversed domain" + ); + let labels = value.split('.').collect::>(); + ensure!( + labels.len() >= 2, + "reversed domain needs at least two labels" + ); + ensure!( + labels.iter().all(|label| { + !label.is_empty() + && label.len() <= 63 + && label + .bytes() + .all(|byte| byte.is_ascii_lowercase() || byte.is_ascii_digit() || byte == b'-') + && label + .as_bytes() + .first() + .is_some_and(u8::is_ascii_alphanumeric) + && label + .as_bytes() + .last() + .is_some_and(u8::is_ascii_alphanumeric) + }), + "invalid reversed domain label" + ); + Ok(labels.into_iter().rev().collect::>().join(".")) +} diff --git a/crates/argand-site-registry/src/adapters/mod.rs b/crates/argand-site-registry/src/adapters/mod.rs index aaa8858..90e5682 100644 --- a/crates/argand-site-registry/src/adapters/mod.rs +++ b/crates/argand-site-registry/src/adapters/mod.rs @@ -4,7 +4,11 @@ use crate::model::{Fact, Format, Record}; use anyhow::ensure; use serde_json::json; -use std::io::{BufRead, Read}; +use std::{ + collections::BTreeSet, + io::{BufRead, Read}, +}; +mod common_crawl_web_graph; pub(crate) mod csv_sources; mod curlie; mod ror; @@ -39,6 +43,7 @@ pub fn adapter( format: Format, maximum_record_bytes: usize, coverage_delta: bool, + web_graph_targets: Option>, ) -> Box { match format { Format::WikidataDump => Box::new(wikidata::Wikidata { @@ -58,6 +63,9 @@ pub fn adapter( Format::RorZip => Box::new(ror::Ror { maximum_record_bytes, }), + Format::CommonCrawlDomainRanksTsv => Box::new(common_crawl_web_graph::DomainRanks { + targets: web_graph_targets.unwrap_or_default(), + }), } } diff --git a/crates/argand-site-registry/src/cli.rs b/crates/argand-site-registry/src/cli.rs index beec8f1..0a21e34 100644 --- a/crates/argand-site-registry/src/cli.rs +++ b/crates/argand-site-registry/src/cli.rs @@ -203,6 +203,11 @@ enum Command { #[arg(long)] maximum_database_growth_bytes: Option, }, + /// Compute the exact candidate-domain scope for a compact Web Graph import. + WebGraphSelection { + #[arg(long)] + database: PathBuf, + }, /// Build a new immutable generation; output must not exist. Build { #[arg(long)] @@ -739,6 +744,9 @@ pub(super) async fn run() -> anyhow::Result<()> { maximum_records, maximum_database_growth_bytes, )?, + Command::WebGraphSelection { database } => serde_json::to_value( + registry::store::web_graph_selection(®istry::store::open(&database)?)?, + )?, Command::ObservationImport { database, generation, diff --git a/crates/argand-site-registry/src/download.rs b/crates/argand-site-registry/src/download.rs index f7f7bc9..3e8e7ae 100644 --- a/crates/argand-site-registry/src/download.rs +++ b/crates/argand-site-registry/src/download.rs @@ -102,6 +102,7 @@ pub fn validate_source_url(source: Source, input: &str) -> anyhow::Result<()> { } Source::Psl => host == "publicsuffix.org" && path == "/list/public_suffix_list.dat", Source::Ror => ror_url(host, path), + Source::CommonCrawlWebGraph => common_crawl_domain_ranks_url(host, path), }; ensure!(allowed, "unreviewed source endpoint: {host}{path}"); Ok(()) @@ -130,6 +131,10 @@ pub async fn download(cache: &Path, request: &Download) -> anyhow::Result 0, "maximum bytes must be positive"); + ensure!( + request.scope != "{candidate_domains}", + "candidate-domain scope token is resolved only by the ordered update command" + ); let requested_proof = requested_checksum_proof(request)?; let key = crate::digest(&serde_json::to_vec(request)?); let dir = cache.join(request.source.key()).join(key); @@ -267,6 +272,25 @@ fn ror_url(host: &str, path: &str) -> bool { && components[5] == "content" } +fn common_crawl_domain_ranks_url(host: &str, path: &str) -> bool { + let components = path.trim_start_matches('/').split('/').collect::>(); + if host != "data.commoncrawl.org" + || components.len() != 5 + || components[0] != "projects" + || components[1] != "hyperlinkgraph" + || components[3] != "domain" + { + return false; + } + let release = components[2]; + release.starts_with("cc-main-") + && release.len() <= 128 + && release + .bytes() + .all(|byte| byte.is_ascii_lowercase() || byte.is_ascii_digit() || byte == b'-') + && components[4] == format!("{release}-domain-ranks.txt.gz") +} + fn requested_checksum_proof(request: &Download) -> anyhow::Result> { match (&request.provider_checksum, &request.provider_checksum_url) { (None, None) => Ok(None), @@ -327,7 +351,11 @@ pub(crate) fn validate_integrity_evidence( && evidence.path().rsplit_once('/').map(|item| item.0) == source_parent && evidence.path().ends_with("/sha256sums.txt") } - Source::Majestic | Source::Crux | Source::Curlie | Source::Psl => false, + Source::Majestic + | Source::Crux + | Source::Curlie + | Source::Psl + | Source::CommonCrawlWebGraph => false, }; ensure!(allowed, "unreviewed provider checksum evidence endpoint"); Ok(()) @@ -536,6 +564,29 @@ fn reserve_psl_refresh(cache: &Path) -> anyhow::Result<()> { mod tests { use super::*; + #[tokio::test] + async fn web_graph_scope_token_requires_ordered_update() -> anyhow::Result<()> { + let root = tempfile::tempdir()?; + let request = Download { + source: Source::CommonCrawlWebGraph, + format: Format::CommonCrawlDomainRanksTsv, + compression: Compression::Gzip, + url: "https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2022-may-jun-aug/domain/cc-main-2022-may-jun-aug-domain-ranks.txt.gz".into(), + snapshot: "cc-main-2022-may-jun-aug".into(), + scope: "{candidate_domains}".into(), + maximum_bytes: 3_000_000_000, + maximum_record_bytes: None, + coverage: None, + provider_checksum: None, + provider_checksum_url: None, + }; + let Some(error) = download(root.path(), &request).await.err() else { + anyhow::bail!("scope token accepted outside update"); + }; + assert!(error.to_string().contains("ordered update command")); + Ok(()) + } + #[test] fn validator_bound_ranges_and_source_allowlist() -> anyhow::Result<()> { let url = "https://downloads.majestic.com/majestic_million.csv"; diff --git a/crates/argand-site-registry/src/model.rs b/crates/argand-site-registry/src/model.rs index 38dff37..c7c6062 100644 --- a/crates/argand-site-registry/src/model.rs +++ b/crates/argand-site-registry/src/model.rs @@ -22,6 +22,8 @@ pub enum Source { Psl, /// Research Organization Registry organization records. Ror, + /// Common Crawl domain-level Web Graph ranks. + CommonCrawlWebGraph, } impl Source { @@ -35,6 +37,7 @@ impl Source { Self::Curlie => "curlie", Self::Psl => "psl", Self::Ror => "ror", + Self::CommonCrawlWebGraph => "common_crawl_web_graph", } } /// Exact SPDX data license. @@ -45,6 +48,7 @@ impl Source { Self::Majestic | Self::Curlie => "CC-BY-3.0", Self::Crux => "CC-BY-4.0", Self::Psl => "MPL-2.0", + Self::CommonCrawlWebGraph => "LicenseRef-Common-Crawl-Terms-of-Use", } } /// Authoritative license evidence page. @@ -57,6 +61,7 @@ impl Source { Self::Curlie => "https://curlie.org/docs/en/license.html", Self::Psl => "https://publicsuffix.org/list/public_suffix_list.dat", Self::Ror => "https://ror.readme.io/docs/data-dump", + Self::CommonCrawlWebGraph => "https://commoncrawl.org/terms-of-use", } } } @@ -79,6 +84,8 @@ pub enum Format { PslText, /// Official ROR release ZIP containing schema 2.1 JSON and CSV. RorZip, + /// Common Crawl's six-column domain-rank TSV. + CommonCrawlDomainRanksTsv, } /// How an immutable source object was checked before import. @@ -408,8 +415,22 @@ impl SourceManifest { | (Source::Curlie, Format::CurlieTarGz) | (Source::Psl, Format::PslText) | (Source::Ror, Format::RorZip) + | ( + Source::CommonCrawlWebGraph, + Format::CommonCrawlDomainRanksTsv + ) ); ensure!(valid, "source/format mismatch"); + if self.source == Source::CommonCrawlWebGraph { + let digest = self + .scope + .strip_prefix("candidate-domains:") + .context("Web Graph scope must bind the candidate-domain digest")?; + ensure!( + valid_digest(digest), + "invalid Web Graph candidate-domain digest" + ); + } if let Some(coverage) = &self.coverage { coverage.validate()?; } diff --git a/crates/argand-site-registry/src/release.rs b/crates/argand-site-registry/src/release.rs index bf81dc5..6f3f485 100644 --- a/crates/argand-site-registry/src/release.rs +++ b/crates/argand-site-registry/src/release.rs @@ -22,6 +22,7 @@ pub fn attribution() -> Value { "curlie":{"license":"CC-BY-3.0","credit":"With content from Curlie.org - the largest human-edited directory of the web. Contribute by submitting a website or becoming an editor.","url":"https://curlie.org/","license_url":"https://creativecommons.org/licenses/by/3.0/","public_display":"Use the prescribed HTML attribution on every page using Curlie content: https://curlie.org/docs/en/license.html"}, "psl":{"license":"MPL-2.0","url":"https://publicsuffix.org/list/","license_url":"https://mozilla.org/MPL/2.0/"}, "ror":{"license":"CC0-1.0","url":"https://ror.org/","license_url":"https://ror.readme.io/docs/data-dump","lineage_note":"ROR location metadata identifies GeoNames as an upstream CC BY 3.0 source","upstream_attribution":{"credit":"GeoNames","url":"https://www.geonames.org/","license_url":"https://creativecommons.org/licenses/by/3.0/"}}, + "common_crawl_web_graph":{"license":"LicenseRef-Common-Crawl-Terms-of-Use","credit":"Common Crawl Foundation Web Graph","url":"https://commoncrawl.org/web-graphs","license_url":"https://commoncrawl.org/terms-of-use","scope":"domain-level harmonic centrality, PageRank, and member-host count; graph rank is not ownership or query popularity"}, "argand_candidate_observer":{"license":"CC0-1.0","url":"https://git.argand.org/nicweyand/argand-site-registry","license_url":"https://creativecommons.org/publicdomain/zero/1.0/","scope":"locally authored observation metadata; captured page content is not redistributed"}, "changes":"Argand normalizes and combines assertions; provider endorsement is not implied."}) } diff --git a/crates/argand-site-registry/src/store.rs b/crates/argand-site-registry/src/store.rs index de3ad52..33e5f01 100644 --- a/crates/argand-site-registry/src/store.rs +++ b/crates/argand-site-registry/src/store.rs @@ -3,13 +3,16 @@ use crate::{ adapters::{self, RecordSink}, - model::{Compression, Record, SourceManifest}, + model::{Compression, Format, Record, SourceManifest}, + normalize::Normalizer, }; use anyhow::{Context, ensure}; use rusqlite::{Connection, OptionalExtension, params}; +use serde::Serialize; use sha2::{Digest, Sha256}; use std::{ cell::RefCell, + collections::BTreeSet, io::{BufReader, Read}, path::Path, rc::Rc, @@ -19,6 +22,59 @@ use std::{ /// Adapter/normalization contract recorded in all generation identities. pub const RULE_VERSION: &str = "argand.site-rules/v4"; +/// Deterministic Web Graph projection selected from already imported website evidence. +#[derive(Clone, Debug, Eq, PartialEq, Serialize)] +pub struct WebGraphSelection { + /// Replacement scope that must be bound into the Web Graph source manifest. + pub scope: String, + /// Number of distinct registrable domains retained from the graph. + pub domains: u64, +} + +/// Computes the exact public-identity domain set used by a compact Web Graph import. +/// +/// # Errors +/// Requires one completed PSL source and at least one valid website assertion. +pub fn web_graph_selection(db: &Connection) -> anyhow::Result { + let (selection, _) = web_graph_targets(db)?; + Ok(selection) +} + +fn web_graph_targets(db: &Connection) -> anyhow::Result<(WebGraphSelection, BTreeSet)> { + let (psl_source, encoded): (String, String) = db + .query_row( + "SELECT f.source_id,f.value FROM facts f JOIN sources s ON s.id=f.source_id WHERE s.complete=1 AND f.predicate='psl' ORDER BY f.source_id LIMIT 1", + [], + |row| Ok((row.get(0)?, row.get(1)?)), + ) + .context("import a complete PSL snapshot before Common Crawl Web Graph")?; + let psl: String = serde_json::from_str(&encoded)?; + let normalizer = Normalizer::new(psl.as_bytes(), psl_source)?; + let mut statement = db.prepare( + "SELECT f.value FROM facts f JOIN sources s ON s.id=f.source_id WHERE s.complete=1 AND f.predicate='website' ORDER BY f.id", + )?; + let values = statement.query_map([], |row| row.get::<_, String>(0))?; + let mut domains = BTreeSet::new(); + for encoded in values { + let value: serde_json::Value = serde_json::from_str(&encoded?)?; + if let Some(url) = value.get("url").and_then(serde_json::Value::as_str) + && let Ok(property) = normalizer.url(url) + { + domains.insert(property.domain.registrable_domain); + } + } + ensure!( + !domains.is_empty(), + "import website assertions before Common Crawl Web Graph" + ); + let digest = crate::digest(&serde_json::to_vec(&domains)?); + let selection = WebGraphSelection { + scope: format!("candidate-domains:{digest}"), + domains: u64::try_from(domains.len())?, + }; + Ok((selection, domains)) +} + /// Whether a signed immutable generation uses a reader-compatible rule contract. #[must_use] pub fn supported_rule_version(version: &str) -> bool { @@ -154,6 +210,17 @@ pub fn import_with_limits( limits: ImportLimits, ) -> anyhow::Result { manifest.validate()?; + let graph_targets = if manifest.format == Format::CommonCrawlDomainRanksTsv { + let (selection, targets) = web_graph_targets(db)?; + ensure!( + manifest.scope == selection.scope, + "Web Graph manifest scope does not match current candidate domains; expected {}", + selection.scope + ); + Some(targets) + } else { + None + }; ensure!( limits.maximum_expanded_bytes > 0 && limits.maximum_records > 0 @@ -225,6 +292,7 @@ pub fn import_with_limits( .coverage .as_ref() .is_some_and(|coverage| coverage.kind == crate::model::CoverageKind::Delta), + graph_targets, ) .ingest(&mut reader, &mut sink); parsed.and_then(|()| { diff --git a/crates/argand-site-registry/src/update.rs b/crates/argand-site-registry/src/update.rs index a24a379..64ca8b5 100644 --- a/crates/argand-site-registry/src/update.rs +++ b/crates/argand-site-registry/src/update.rs @@ -58,6 +58,7 @@ pub async fn run(config: &Config) -> anyhow::Result { .open(config.generations.join("update.lock"))?; lock.try_lock().context("registry update already running")?; let mut inputs = config.inputs.clone(); + let mut graph_downloads = Vec::new(); let mut db = crate::store::open(&config.database)?; let now = Utc::now(); for request in &config.downloads { @@ -66,6 +67,10 @@ pub async fn run(config: &Config) -> anyhow::Result { .snapshot .replace("{date}", &now.format("%Y-%m-%d").to_string()) .replace("{month}", &now.format("%Y-%m").to_string()); + if request.source == crate::model::Source::CommonCrawlWebGraph { + graph_downloads.push(request); + continue; + } if config.auto_supersede_typed_snapshots && let Some(coverage) = &mut request.coverage && matches!( @@ -119,8 +124,31 @@ pub async fn run(config: &Config) -> anyhow::Result { } inputs.push(crate::crux::download(&config.cache, &request).await?); } - anyhow::ensure!(!inputs.is_empty(), "update config contains no sources"); + let mut graph_inputs = Vec::new(); + let mut identity_inputs = Vec::new(); for input in inputs { + let manifest: SourceManifest = crate::read_json(&input.manifest)?; + if manifest.source == crate::model::Source::CommonCrawlWebGraph { + graph_inputs.push(input); + } else { + identity_inputs.push(input); + } + } + anyhow::ensure!( + !identity_inputs.is_empty() || !graph_inputs.is_empty() || !graph_downloads.is_empty(), + "update config contains no sources" + ); + for input in identity_inputs { + let manifest: SourceManifest = crate::read_json(&input.manifest)?; + crate::store::import(&mut db, &manifest, &input.input)?; + } + for mut request in graph_downloads { + if request.scope == "{candidate_domains}" { + request.scope = crate::store::web_graph_selection(&db)?.scope; + } + graph_inputs.push(crate::download::download(&config.cache, &request).await?); + } + for input in graph_inputs { let manifest: SourceManifest = crate::read_json(&input.manifest)?; crate::store::import(&mut db, &manifest, &input.input)?; } diff --git a/crates/argand-site-registry/tests/common/mod.rs b/crates/argand-site-registry/tests/common/mod.rs index 8222f08..1e50244 100644 --- a/crates/argand-site-registry/tests/common/mod.rs +++ b/crates/argand-site-registry/tests/common/mod.rs @@ -71,6 +71,9 @@ pub fn manifest(source: Source, format: Format, bytes: &[u8]) -> anyhow::Result< Source::Ror => { "https://zenodo.org/api/records/22099990/files/v2.12-2026-08-25-ror-data.zip/content" } + Source::CommonCrawlWebGraph => { + "https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2022-may-jun-aug/domain/cc-main-2022-may-jun-aug-domain-ranks.txt.gz" + } }; Ok(SourceManifest { schema: "argand.site-source/v1".into(), diff --git a/crates/argand-site-registry/tests/webgraph.rs b/crates/argand-site-registry/tests/webgraph.rs new file mode 100644 index 0000000..6f3e8a8 --- /dev/null +++ b/crates/argand-site-registry/tests/webgraph.rs @@ -0,0 +1,239 @@ +// By Nic Weyand! +//! Common Crawl Web Graph evidence stays source separated and cannot authorize routes. + +#![allow(dead_code)] // Shared integration helpers intentionally cover a wider fixture surface. + +mod common; + +use argand_site_registry::{ + download::validate_source_url, + model::{Compression, Format, Source, SourceManifest}, + query::ResolutionStatus, + store, +}; +use std::{fmt::Write as _, time::Instant}; + +const WEBGRAPH: &str = "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n\ +1\t3.2914686E7\t1\t0.018076941061056315\tcom.googleapis\t4482\n\ +2\t3.2131562E7\t3\t0.012273178013351222\tcom.facebook\t18795\n"; + +fn import_identity_candidates( + db: &mut rusqlite::Connection, + root: &std::path::Path, +) -> anyhow::Result<()> { + common::import( + db, + root, + Source::Psl, + Format::PslText, + common::PSL.as_bytes(), + )?; + common::import( + db, + root, + Source::Wikidata, + Format::WikidataEntities, + &serde_json::to_vec(&common::wikidata())?, + )?; + Ok(()) +} + +fn graph_manifest(db: &rusqlite::Connection, bytes: &[u8]) -> anyhow::Result { + let mut manifest = common::manifest( + Source::CommonCrawlWebGraph, + Format::CommonCrawlDomainRanksTsv, + bytes, + )?; + manifest.scope = store::web_graph_selection(db)?.scope; + Ok(manifest) +} + +fn import_graph( + db: &mut rusqlite::Connection, + root: &std::path::Path, + bytes: &[u8], +) -> anyhow::Result<()> { + let manifest = graph_manifest(db, bytes)?; + let input = root.join(format!("{}.graph", manifest.sha256)); + std::fs::write(&input, bytes)?; + store::import(db, &manifest, &input)?; + Ok(()) +} + +#[test] +fn domain_ranks_are_popularity_only_and_preserve_provider_fields() -> anyhow::Result<()> { + let root = tempfile::tempdir()?; + let mut db = store::open(&root.path().join("store.sqlite"))?; + import_identity_candidates(&mut db, root.path())?; + let identity_before: i64 = db.query_row( + "SELECT count(*) FROM facts WHERE predicate NOT IN ('psl','popularity')", + [], + |row| row.get(0), + )?; + import_graph(&mut db, root.path(), WEBGRAPH.as_bytes())?; + + let identity_facts: i64 = db.query_row( + "SELECT count(*) FROM facts WHERE predicate NOT IN ('psl','popularity')", + [], + |row| row.get(0), + )?; + assert_eq!(identity_facts, identity_before); + for table in ["reviews", "votes"] { + let count: i64 = db.query_row(&format!("SELECT count(*) FROM {table}"), [], |row| { + row.get(0) + })?; + assert_eq!(count, 0, "Web Graph unexpectedly populated {table}"); + } + + let first = common::build(&db, root.path(), "first")?; + let second = common::build(&db, root.path(), "second")?; + assert_eq!(first.identity, second.identity); + assert_eq!(first.lookup("Facebook", 10)?.total_entities, 1); + assert_ne!( + first + .resolve_explained("Facebook", None, None, common::timestamp()?)? + .status, + ResolutionStatus::Resolved + ); + + assert_eq!(first.popularity("googleapis.com", 10)?.total, 0); + + let lookup = first.popularity("facebook.com", 10)?; + assert_eq!(lookup.total, 1); + let observation = &lookup.observations[0]; + assert_eq!(observation.source, "common_crawl_web_graph"); + assert_eq!(observation.target, "facebook.com"); + assert_eq!(observation.value["harmonic_rank"], 2); + assert_eq!(observation.value["harmonic_value"], 3.213_156_2E7); + assert_eq!(observation.value["pagerank_rank"], 3); + assert_eq!( + observation.value["pagerank_value"], + 0.012_273_178_013_351_222_f64 + ); + assert_eq!(observation.value["member_hosts"], 18_795); + assert!(observation.value.get("source_hosts").is_none()); + assert_eq!( + observation.provenance["source"]["license"], + "LicenseRef-Common-Crawl-Terms-of-Use" + ); + let native_id: String = db.query_row( + "SELECT r.native_id FROM records r JOIN sources s ON s.id=r.source_id WHERE s.source='common_crawl_web_graph'", + [], + |row| row.get(0), + )?; + assert_eq!(native_id, "row:2"); + Ok(()) +} + +#[test] +fn graph_import_requires_identity_candidates_and_exact_bound_scope() -> anyhow::Result<()> { + let root = tempfile::tempdir()?; + let mut db = store::open(&root.path().join("store.sqlite"))?; + let mut manifest = common::manifest( + Source::CommonCrawlWebGraph, + Format::CommonCrawlDomainRanksTsv, + WEBGRAPH.as_bytes(), + )?; + manifest.scope = format!("candidate-domains:{}", "0".repeat(64)); + let input = root.path().join("graph.tsv"); + std::fs::write(&input, WEBGRAPH)?; + assert!(store::import(&mut db, &manifest, &input).is_err()); + + import_identity_candidates(&mut db, root.path())?; + assert!(store::import(&mut db, &manifest, &input).is_err()); + manifest.scope = store::web_graph_selection(&db)?.scope; + store::import(&mut db, &manifest, &input)?; + Ok(()) +} + +#[test] +fn domain_rank_parser_rejects_schema_and_value_drift() -> anyhow::Result<()> { + for bad in [ + "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\n", + "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n\n", + "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n0\t1\t1\t1\tcom.example\t1\n", + "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n1\tNaN\t1\t1\tcom.example\t1\n", + "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n1\t1\t1\t1\tcom..example\t1\n", + "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n1\t1\t1\t1\tcom.example\t0\n", + ] { + let root = tempfile::tempdir()?; + let mut db = store::open(&root.path().join("store.sqlite"))?; + import_identity_candidates(&mut db, root.path())?; + assert!( + import_graph(&mut db, root.path(), bad.as_bytes()).is_err(), + "accepted malformed Web Graph input {bad:?}" + ); + } + Ok(()) +} + +#[test] +fn domain_rank_source_url_is_exactly_allowlisted() -> anyhow::Result<()> { + let good = "https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2022-may-jun-aug/domain/cc-main-2022-may-jun-aug-domain-ranks.txt.gz"; + validate_source_url(Source::CommonCrawlWebGraph, good)?; + for bad in [ + "http://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2022-may-jun-aug/domain/cc-main-2022-may-jun-aug-domain-ranks.txt.gz", + "https://data.commoncrawl.org.evil.example/projects/hyperlinkgraph/cc-main-2022-may-jun-aug/domain/cc-main-2022-may-jun-aug-domain-ranks.txt.gz", + "https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2022-may-jun-aug/host/cc-main-2022-may-jun-aug-host-ranks.txt.gz", + "https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2022-may-jun-aug/domain/other-domain-ranks.txt.gz", + "https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2022-may-jun-aug/domain/cc-main-2022-may-jun-aug-domain-ranks.txt.gz?x=1", + ] { + assert!( + validate_source_url(Source::CommonCrawlWebGraph, bad).is_err(), + "accepted unreviewed Web Graph URL {bad}" + ); + } + Ok(()) +} + +#[test] +fn gzip_source_is_consumed_and_authenticated_end_to_end() -> anyhow::Result<()> { + let root = tempfile::tempdir()?; + let compressed = common::gzip(WEBGRAPH.as_bytes())?; + let input = root.path().join("domain-ranks.txt.gz"); + std::fs::write(&input, &compressed)?; + let mut db = store::open(&root.path().join("store.sqlite"))?; + import_identity_candidates(&mut db, root.path())?; + let mut manifest = graph_manifest(&db, &compressed)?; + manifest.compression = Compression::Gzip; + store::import(&mut db, &manifest, &input)?; + let registry = common::build(&db, root.path(), "gzip")?; + assert_eq!(registry.popularity("facebook.com", 10)?.total, 1); + Ok(()) +} + +#[test] +#[ignore = "explicit resource benchmark"] +fn hundred_thousand_rows_import_within_engineering_budget() -> anyhow::Result<()> { + let mut input = + String::from("#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n"); + for rank in 1..=100_000_u32 { + let reversed_domain = if rank == 50_000 { + "com.facebook".to_owned() + } else { + format!("com.example-{rank}") + }; + writeln!( + input, + "{rank}\t{}\t{rank}\t{}\t{reversed_domain}\t1", + 100_001 - rank, + 0.1 + )?; + } + let root = tempfile::tempdir()?; + let mut db = store::open(&root.path().join("store.sqlite"))?; + import_identity_candidates(&mut db, root.path())?; + let started = Instant::now(); + import_graph(&mut db, root.path(), input.as_bytes())?; + assert!( + started.elapsed().as_secs_f64() < 10.0, + "100k Web Graph rows exceeded the 10-second engineering budget" + ); + let retained: i64 = db.query_row( + "SELECT count(*) FROM facts WHERE predicate='popularity'", + [], + |row| row.get(0), + )?; + assert_eq!(retained, 1); + Ok(()) +} diff --git a/docs/CONSUMERS.md b/docs/CONSUMERS.md index 20ddad9..c072268 100644 --- a/docs/CONSUMERS.md +++ b/docs/CONSUMERS.md @@ -34,7 +34,7 @@ source attribution and application-specific malware/content policy. ## Current contracts -Code version 0.5.0 uses writer schema 5 and `argand.site-rules/v4`. +Code version 0.6.0 uses writer schema 5 and `argand.site-rules/v4`. New compact `COMPLETE.json` files use `argand.site-registry/v3` and bind: - authenticated `registry.sqlite` bytes; @@ -101,8 +101,9 @@ revocation history. Never mutate a complete generation to migrate it. See ## Argand integration `UPSTREAM.json` records the original Argand extraction baseline and file hashes. -Argand pins signed v0.5.0 revision -`3d3e08cdfd303df9fbd347a9bab2ba52ad575759`. The public beta uses Site Registry as +Argand's prior integration pinned signed v0.5.0 revision +`3d3e08cdfd303df9fbd347a9bab2ba52ad575759`; the v0.6 downstream pin is recorded +by the Argand integration commit after this source release. The public beta uses Site Registry as Navigate's authoritative auto-route catalog. Its native `navigation-catalog/v2` file is only a collection- and content-policy-bound serving projection compiled from one exact registry generation; it is not a second independently curated diff --git a/docs/FORMATS.md b/docs/FORMATS.md index af0c1e6..e87cbf2 100644 --- a/docs/FORMATS.md +++ b/docs/FORMATS.md @@ -1,8 +1,18 @@ # Versioned formats -Version 0.5 uses writer schema 5 and `argand.site-rules/v4`. Schema identifiers +Version 0.6 uses writer schema 5 and `argand.site-rules/v4`. Schema identifiers are independent from the crate version. Unknown schemas and rules fail closed. +The Common Crawl domain-rank adapter is new in 0.6, so no older v4 store can +contain one of its source manifests. Its replacement scope is +`candidate-domains:`, where the digest covers the canonical JSON encoding +of the sorted set of registrable domains derived from retained complete website +assertions. Retained superseded evidence may enlarge this conservative set but +cannot create a route or make a retired assertion active. +The importer authenticates and validates every graph row but persists only the +selected domains. A changed identity frontier therefore produces a new immutable +source identity instead of silently reusing a stale projection. + | Artifact | Current schema | Purpose | | --- | --- | --- | | Source manifest | `argand.site-source/v3` | Exact source object, integrity proof, lineage, parser bound and typed coverage | diff --git a/docs/SOURCE-CANDIDATES.md b/docs/SOURCE-CANDIDATES.md index 1301f54..54487b4 100644 --- a/docs/SOURCE-CANDIDATES.md +++ b/docs/SOURCE-CANDIDATES.md @@ -7,13 +7,14 @@ all been reviewed. A research entry below is not permission to ingest it. | Source | Status | Decision and next gate | | --- | --- | --- | | ROR | Admitted in 0.5 | The official CC0 schema 2.1 ZIP is streamed with exact Zenodo checksum evidence. Organization websites remain assertions; inactive/withdrawn edges are ineligible. GeoNames location lineage is explicit. | +| Common Crawl domain Web Graph ranks | Admitted after 0.5 as authority evidence | The official six-column domain-rank object is streamed from an exact allowlisted release URL under Common Crawl's Terms of Use. Harmonic centrality, PageRank and member-host count remain source separated. The source cannot create identities, official-site edges, reviews, or routes. Full-graph acquisition and production-catalog selection remain separate operational gates. | | MusicBrainz | Next adapter; held | The [official download documentation](https://musicbrainz.org/doc/MusicBrainz_Database/Download) identifies the core `mbdump.tar.bz2` snapshot as CC0. The live replication/edit/statistics material with noncommercial terms is excluded. Admission still needs a current core snapshot/checksum canary and a bounded relational-table adapter for documented [URL relationships](https://musicbrainz.org/doc/Style/Relationships/URLs). | | GND | Research hold | The [DNB open-data distribution](https://data.dnb.de/opendata/) must be checked at implementation time for the exact file license, current JSON-LD/RDF predicates, checksum and useful homepage coverage. Stop the adapter if explicit homepage coverage does not justify it. | | ORCID public data | Research hold | Its self-declared links need an individuals-only privacy, impersonation and volatility policy in addition to the [public-file terms](https://info.orcid.org/public-data-file-use-policy/). It could never auto-approve a route. | | OpenAlex institutions | Correlated-source hold | Institution metadata can inherit ROR. Any future use must declare ROR upstream and cannot count as independent website corroboration. See the [institution source documentation](https://help.openalex.org/data/institutions/). | | OpenStreetMap | License-architecture hold | No ingestion until an ODbL-compatible attribution, database-right and redistribution design is accepted. See the [OSMF license FAQ](https://osmfoundation.org/wiki/Licence_and_Legal_FAQ). | | Government/corporate registries | Jurisdiction hold | Review one jurisdiction and exact field at a time. Stable identifiers may support crosswalks; the registry cannot infer a website absent an authoritative field. | -| DNS, RDAP, certificate transparency, package registries, web crawl data | Observation-only research | Exact commercial reuse terms and retention rules must be approved first. These sources describe current infrastructure and cannot establish entity ownership alone. | +| DNS, RDAP, certificate transparency, package registries, other web crawl data | Observation-only research | Exact commercial reuse terms and retention rules must be approved first. These sources describe current infrastructure and cannot establish entity ownership alone. | | Open Library and unresolved-rights sources | Excluded | Keep excluded until the underlying data rights and redistribution obligations are clear enough for commercial reuse. | Cloudflare Radar, default Tranco, Cisco Umbrella, arbitrary mirrors, and any diff --git a/docs/VALIDATION.md b/docs/VALIDATION.md index 8f6afb3..eccb950 100644 --- a/docs/VALIDATION.md +++ b/docs/VALIDATION.md @@ -1,5 +1,27 @@ # Validation +# Version 0.6.0 Web Graph and updater validation, 2026-09-22 + +The evaluation contract was written before implementation. The new Common Crawl +domain-rank integration passed five default integration tests; its sixth test is +an explicit resource benchmark. The benchmark parsed 100,000 valid provider-shaped +rows, retained exactly one identity-matched domain, and completed in 0.52 seconds +of test time. The enclosing warm Cargo process used 79,944 KiB peak RSS, wrote +9,136 filesystem blocks, and used no swap on the development machine. These are +engineering bounds, not a full 2 GiB provider-object throughput claim. + +`scripts/check.sh` passed after the v0.6 version and Rustls lockfile updates. It +covered formatting, offline all-target checks, strict Clippy, all Rust and CLI +tests, documentation, Python release-package tests, and native/Python consumer +parity. The focused graph suite proves full-stream schema/value validation, +source-line provenance, gzip authentication, exact URL allowlisting, deterministic +builds, compact candidate-domain selection, and zero route authorization from +rank evidence. + +The networked `cargo audit --deny warnings` gate initially detected +RUSTSEC-2026-0285 in Rustls 0.23.43. The lockfile was updated to Rustls 0.23.45; +the repeated audit passed with no findings. + ## Version 0.5.0 release and security validation, 2026-09-13 Implementation commit: `557ba7cd6982b02754d34fb99cba5a116f78f153`, signed by From 3b6966c81a0b4de31cc18f49c39217ae107e6323 Mon Sep 17 00:00:00 2001 From: nicweyand Date: Tue, 22 Sep 2026 09:01:39 -0400 Subject: [PATCH 3/5] fix: accept non-DNS Common Crawl graph rows --- CHANGELOG.md | 8 +++++ Cargo.lock | 4 +-- Cargo.toml | 2 +- .../src/adapters/common_crawl_web_graph.rs | 34 ++++++++++--------- crates/argand-site-registry/tests/webgraph.rs | 25 +++++++++++++- docs/CONSUMERS.md | 2 +- docs/VALIDATION.md | 6 ++++ 7 files changed, 60 insertions(+), 21 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 8d83e1a..ecd31f9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,13 @@ # Changelog +## 0.6.1 - 2026-09-22 + +- Accept complete Common Crawl domain-rank releases containing provider rows + that are not valid DNS hostnames. Such rows remain authenticated and counted + in source coordinates but cannot match or enter the official-domain catalog. +- Keep malformed graph schemas and numeric fields fail-closed, with a regression + derived from the real `com.your_domain` provider row. + ## 0.6.0 - 2026-09-22 - Add a streaming Common Crawl domain Web Graph adapter for harmonic-centrality, diff --git a/Cargo.lock b/Cargo.lock index afcb5e2..1ec3b30 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -78,14 +78,14 @@ checksum = "330a5ed07fa54e4702c9d6c4174f74427fc0ef6e214bbd677ae50a5099946470" [[package]] name = "argand-atomic" -version = "0.6.0" +version = "0.6.1" dependencies = [ "tempfile", ] [[package]] name = "argand-site-registry" -version = "0.6.0" +version = "0.6.1" dependencies = [ "anyhow", "argand-atomic", diff --git a/Cargo.toml b/Cargo.toml index 6bfc5f0..a4f7ba4 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -4,7 +4,7 @@ resolver = "3" members = ["crates/argand-atomic", "crates/argand-site-registry"] [workspace.package] -version = "0.6.0" +version = "0.6.1" authors = ["Nic Weyand"] edition = "2024" license = "AGPL-3.0-or-later" diff --git a/crates/argand-site-registry/src/adapters/common_crawl_web_graph.rs b/crates/argand-site-registry/src/adapters/common_crawl_web_graph.rs index 71c835c..96d8505 100644 --- a/crates/argand-site-registry/src/adapters/common_crawl_web_graph.rs +++ b/crates/argand-site-registry/src/adapters/common_crawl_web_graph.rs @@ -43,9 +43,15 @@ impl SourceAdapter for DomainRanks { let harmonic_value = nonnegative_finite(fields[1], "harmonic value")?; let pagerank_rank = positive_integer(fields[2], "PageRank rank")?; let pagerank_value = nonnegative_finite(fields[3], "PageRank value")?; - let target = reverse_domain(fields[4])?; let member_hosts = positive_integer(fields[5], "member host count")?; source_row += 1; + let Some(target) = reverse_domain(fields[4]) else { + // The provider graph contains a small amount of underscore and + // otherwise non-DNS host material. It cannot match the registry's + // normalized public identity domains, but it remains part of the + // authenticated input stream and source-row coordinate space. + continue; + }; if !self.targets.contains(&target) { continue; } @@ -118,18 +124,13 @@ fn nonnegative_finite(value: &str, field: &str) -> anyhow::Result { Ok(parsed) } -fn reverse_domain(value: &str) -> anyhow::Result { - ensure!( - !value.is_empty() && value.len() <= 253 && value == value.trim(), - "invalid reversed domain" - ); +fn reverse_domain(value: &str) -> Option { + if value.is_empty() || value.len() > 253 || value != value.trim() { + return None; + } let labels = value.split('.').collect::>(); - ensure!( - labels.len() >= 2, - "reversed domain needs at least two labels" - ); - ensure!( - labels.iter().all(|label| { + if labels.len() < 2 + || !labels.iter().all(|label| { !label.is_empty() && label.len() <= 63 && label @@ -143,8 +144,9 @@ fn reverse_domain(value: &str) -> anyhow::Result { .as_bytes() .last() .is_some_and(u8::is_ascii_alphanumeric) - }), - "invalid reversed domain label" - ); - Ok(labels.into_iter().rev().collect::>().join(".")) + }) + { + return None; + } + Some(labels.into_iter().rev().collect::>().join(".")) } diff --git a/crates/argand-site-registry/tests/webgraph.rs b/crates/argand-site-registry/tests/webgraph.rs index 6f3e8a8..f827ca1 100644 --- a/crates/argand-site-registry/tests/webgraph.rs +++ b/crates/argand-site-registry/tests/webgraph.rs @@ -153,7 +153,6 @@ fn domain_rank_parser_rejects_schema_and_value_drift() -> anyhow::Result<()> { "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n\n", "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n0\t1\t1\t1\tcom.example\t1\n", "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n1\tNaN\t1\t1\tcom.example\t1\n", - "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n1\t1\t1\t1\tcom..example\t1\n", "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n1\t1\t1\t1\tcom.example\t0\n", ] { let root = tempfile::tempdir()?; @@ -167,6 +166,30 @@ fn domain_rank_parser_rejects_schema_and_value_drift() -> anyhow::Result<()> { Ok(()) } +#[test] +fn non_dns_provider_rows_are_skipped_without_losing_source_coordinates() -> anyhow::Result<()> { + let root = tempfile::tempdir()?; + let mut db = store::open(&root.path().join("store.sqlite"))?; + import_identity_candidates(&mut db, root.path())?; + let input = "#harmonicc_pos\t#harmonicc_val\t#pr_pos\t#pr_val\t#host_rev\t#n_hosts\n\ +1\t1\t1\t1\tcom.your_domain\t15\n\ +2\t1\t2\t1\tcom.facebook\t18795\n"; + import_graph(&mut db, root.path(), input.as_bytes())?; + + let retained: Vec<(String, String)> = { + let mut statement = db.prepare( + "SELECT r.native_id,f.value FROM records r JOIN facts f ON f.source_id=r.source_id AND f.ordinal=r.ordinal WHERE f.predicate='popularity' ORDER BY r.native_id", + )?; + statement + .query_map([], |row| Ok((row.get(0)?, row.get(1)?)))? + .collect::>()? + }; + assert_eq!(retained.len(), 1); + assert_eq!(retained[0].0, "row:2"); + assert!(retained[0].1.contains("facebook.com")); + Ok(()) +} + #[test] fn domain_rank_source_url_is_exactly_allowlisted() -> anyhow::Result<()> { let good = "https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2022-may-jun-aug/domain/cc-main-2022-may-jun-aug-domain-ranks.txt.gz"; diff --git a/docs/CONSUMERS.md b/docs/CONSUMERS.md index c072268..1433fdb 100644 --- a/docs/CONSUMERS.md +++ b/docs/CONSUMERS.md @@ -34,7 +34,7 @@ source attribution and application-specific malware/content policy. ## Current contracts -Code version 0.6.0 uses writer schema 5 and `argand.site-rules/v4`. +Code version 0.6.1 uses writer schema 5 and `argand.site-rules/v4`. New compact `COMPLETE.json` files use `argand.site-registry/v3` and bind: - authenticated `registry.sqlite` bytes; diff --git a/docs/VALIDATION.md b/docs/VALIDATION.md index eccb950..65e9b7b 100644 --- a/docs/VALIDATION.md +++ b/docs/VALIDATION.md @@ -2,6 +2,12 @@ # Version 0.6.0 Web Graph and updater validation, 2026-09-22 +Patch release 0.6.1 additionally replays the real provider-shaped +`com.your_domain` case: the row remains in authenticated input/coordinate +accounting but is not retained as a DNS target. The following valid +`com.facebook` row is retained at its original `row:2` coordinate. Schema and +numeric corruption continue to fail closed. + The evaluation contract was written before implementation. The new Common Crawl domain-rank integration passed five default integration tests; its sixth test is an explicit resource benchmark. The benchmark parsed 100,000 valid provider-shaped From 3c89a540d910cb298ba752880efeb1ed157dab43 Mon Sep 17 00:00:00 2001 From: nicweyand Date: Tue, 22 Sep 2026 09:09:41 -0400 Subject: [PATCH 4/5] fix: retain PSL in compact catalogs --- CHANGELOG.md | 5 +++++ Cargo.lock | 4 ++-- Cargo.toml | 2 +- crates/argand-site-registry/src/audit.rs | 3 ++- crates/argand-site-registry/tests/v05.rs | 3 +++ docs/CONSUMERS.md | 2 +- docs/VALIDATION.md | 2 ++ 7 files changed, 16 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index ecd31f9..a3bea2f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,10 @@ # Changelog +## 0.6.2 - 2026-09-22 + +- Retain the PSL normalization fact in compact catalogs and exercise the review + queue against the compact runtime rather than only the full writer projection. + ## 0.6.1 - 2026-09-22 - Accept complete Common Crawl domain-rank releases containing provider rows diff --git a/Cargo.lock b/Cargo.lock index 1ec3b30..506a25f 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -78,14 +78,14 @@ checksum = "330a5ed07fa54e4702c9d6c4174f74427fc0ef6e214bbd677ae50a5099946470" [[package]] name = "argand-atomic" -version = "0.6.1" +version = "0.6.2" dependencies = [ "tempfile", ] [[package]] name = "argand-site-registry" -version = "0.6.1" +version = "0.6.2" dependencies = [ "anyhow", "argand-atomic", diff --git a/Cargo.toml b/Cargo.toml index a4f7ba4..6f02cbe 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -4,7 +4,7 @@ resolver = "3" members = ["crates/argand-atomic", "crates/argand-site-registry"] [workspace.package] -version = "0.6.1" +version = "0.6.2" authors = ["Nic Weyand"] edition = "2024" license = "AGPL-3.0-or-later" diff --git a/crates/argand-site-registry/src/audit.rs b/crates/argand-site-registry/src/audit.rs index a9d270f..392b308 100644 --- a/crates/argand-site-registry/src/audit.rs +++ b/crates/argand-site-registry/src/audit.rs @@ -616,7 +616,8 @@ pub(crate) fn compact_runtime(db: &Connection) -> anyhow::Result<()> { INSERT OR IGNORE INTO runtime_facts SELECT fact FROM names; INSERT OR IGNORE INTO runtime_facts SELECT fact FROM popularity; INSERT OR IGNORE INTO runtime_facts SELECT fact FROM rejected; - INSERT OR IGNORE INTO runtime_facts SELECT item.value FROM edges,json_each(edges.facts) item;", + INSERT OR IGNORE INTO runtime_facts SELECT item.value FROM edges,json_each(edges.facts) item; + INSERT OR IGNORE INTO runtime_facts SELECT id FROM facts WHERE predicate='psl';", )?; db.execute( "DELETE FROM facts WHERE NOT EXISTS(SELECT 1 FROM runtime_facts r WHERE r.id=facts.id)", diff --git a/crates/argand-site-registry/tests/v05.rs b/crates/argand-site-registry/tests/v05.rs index 793e231..023abee 100644 --- a/crates/argand-site-registry/tests/v05.rs +++ b/crates/argand-site-registry/tests/v05.rs @@ -267,6 +267,9 @@ fn compact_generation_keeps_runtime_results_and_authenticates_cold_history() -> serde_json::to_value(full.lookup("Facebook", 20)?.candidates)?, serde_json::to_value(compact.lookup("Facebook", 20)?.candidates)? ); + let compact_queue = + argand_site_registry::queue::review_queue(&compact, common::timestamp()?, 100, 1_000)?; + assert!(!compact_queue.items.is_empty()); assert!( std::fs::metadata(compact_path.join("registry.sqlite"))?.len() < std::fs::metadata(root.path().join("full/registry.sqlite"))?.len() diff --git a/docs/CONSUMERS.md b/docs/CONSUMERS.md index 1433fdb..b17c1f4 100644 --- a/docs/CONSUMERS.md +++ b/docs/CONSUMERS.md @@ -34,7 +34,7 @@ source attribution and application-specific malware/content policy. ## Current contracts -Code version 0.6.1 uses writer schema 5 and `argand.site-rules/v4`. +Code version 0.6.2 uses writer schema 5 and `argand.site-rules/v4`. New compact `COMPLETE.json` files use `argand.site-registry/v3` and bind: - authenticated `registry.sqlite` bytes; diff --git a/docs/VALIDATION.md b/docs/VALIDATION.md index 65e9b7b..309ad3f 100644 --- a/docs/VALIDATION.md +++ b/docs/VALIDATION.md @@ -7,6 +7,8 @@ Patch release 0.6.1 additionally replays the real provider-shaped accounting but is not retained as a DNS target. The following valid `com.facebook` row is retained at its original `row:2` coordinate. Schema and numeric corruption continue to fail closed. +The compact-generation lifecycle now also executes the review queue, proving its +PSL normalization input remains in the runtime catalog. The evaluation contract was written before implementation. The new Common Crawl domain-rank integration passed five default integration tests; its sixth test is From 62e7a67cba70cffd4672102b064aceecac857ca7 Mon Sep 17 00:00:00 2001 From: nicweyand Date: Tue, 22 Sep 2026 09:29:15 -0400 Subject: [PATCH 5/5] docs: publish the v0.6.2 catalog --- README.md | 2 +- docs/CONSUMERS.md | 8 +++--- docs/PUBLIC_CATALOG.md | 56 ++++++++++++++++++++++++++++-------------- 3 files changed, 44 insertions(+), 22 deletions(-) diff --git a/README.md b/README.md index 93b18eb..c225623 100644 --- a/README.md +++ b/README.md @@ -15,7 +15,7 @@ facebook -> Facebook (Wikidata Q355) -> https://www.facebook.com/ The repository contains the library, CLI, schemas, migrations, synthetic fixtures, and independently authenticated public trust roots. It does not place a mutable production database in Git. Publishers import source evidence, collect -signed reviews, and distribute immutable signed registry generations. The first +signed reviews, and distribute immutable signed registry generations. The current public catalog generation is available as a release asset; see [Public catalog](docs/PUBLIC_CATALOG.md). diff --git a/docs/CONSUMERS.md b/docs/CONSUMERS.md index b17c1f4..f108cdc 100644 --- a/docs/CONSUMERS.md +++ b/docs/CONSUMERS.md @@ -102,9 +102,11 @@ revocation history. Never mutate a complete generation to migrate it. See `UPSTREAM.json` records the original Argand extraction baseline and file hashes. Argand's prior integration pinned signed v0.5.0 revision -`3d3e08cdfd303df9fbd347a9bab2ba52ad575759`; the v0.6 downstream pin is recorded -by the Argand integration commit after this source release. The public beta uses Site Registry as -Navigate's authoritative auto-route catalog. Its native `navigation-catalog/v2` +`3d3e08cdfd303df9fbd347a9bab2ba52ad575759`. Argand now pins the public signed +v0.6.2 source release at `3c89a540d910cb298ba752880efeb1ed157dab43` in +downstream commit `224616eb9f6685d1a656b113b8460fb80c0c5a6b`. +The public beta uses Site Registry as Navigate's authoritative auto-route +catalog. Its native `navigation-catalog/v2` file is only a collection- and content-policy-bound serving projection compiled from one exact registry generation; it is not a second independently curated destination catalog. diff --git a/docs/PUBLIC_CATALOG.md b/docs/PUBLIC_CATALOG.md index dbbdf89..c723bee 100644 --- a/docs/PUBLIC_CATALOG.md +++ b/docs/PUBLIC_CATALOG.md @@ -1,14 +1,14 @@ # Public signed catalog -The v0.5.0 Forgejo release publishes the first immutable data generation that any -Site Registry consumer can verify and resolve: +The v0.6.2 Forgejo release publishes the current immutable data generation that +Argand and any other Site Registry consumer can verify and resolve: -- release: -- asset: `argand-site-registry-catalog-20260920-v1.tar.gz` +- release: +- asset: `argand-site-registry-catalog-v0.6.2.tar.gz` - asset SHA-256: - `d878fa057397effa5dc729d2fa3a689c8edd1f4112ef1326dd6131b3fdeab63e` + `48b0cdf453862d858c4bec6c564360e1309605e30af9aba1f54a9446b9bdbe41` - generation pin: - `ede14746da8817aafdf705dd88cfeabbe8d23e1991e43a304acd8eca9249b18a` + `5e5d8fd5dc1864dc3f4c53ec71cb5ac64f6db592cfbc8cc56f48a444378e2309` The release also carries a checksum file and an OpenSSH signature under namespace `argand-site-registry-release`. Verify it against @@ -17,13 +17,13 @@ The signed Git history is the independent channel for the trust root; do not lea the only trusted key from the archive it authenticates. ```bash -sha256sum --check argand-site-registry-catalog-20260920-v1.tar.gz.sha256 +sha256sum --check argand-site-registry-catalog-v0.6.2.tar.gz.sha256 ssh-keygen -Y verify \ -f trust/public-catalog-20260920/publisher-allowed-signers \ -I argand-site-registry-publisher-v1 \ -n argand-site-registry-release \ - -s argand-site-registry-catalog-20260920-v1.tar.gz.sig \ - < argand-site-registry-catalog-20260920-v1.tar.gz + -s argand-site-registry-catalog-v0.6.2.tar.gz.sig \ + < argand-site-registry-catalog-v0.6.2.tar.gz ``` After extraction, verify every member with `SHA256SUMS`, then authenticate the @@ -31,25 +31,36 @@ generation and exact reviewer trust root: ```bash argand-site-registry activate \ - --generation public-release-v0.5.0/catalog \ + --generation public-release-v0.6.2/catalog \ --current current.json \ --allowed-signers trust/public-catalog-20260920/publisher-allowed-signers \ --allowed-reviewers trust/public-catalog-20260920/reviewer-allowed-signers \ --identity argand-site-registry-publisher-v1 argand-site-registry resolve \ - --generation public-release-v0.5.0/catalog \ - --pin ede14746da8817aafdf705dd88cfeabbe8d23e1991e43a304acd8eca9249b18a \ - --query "yahoo mail" + --generation public-release-v0.6.2/catalog \ + --pin 5e5d8fd5dc1864dc3f4c53ec71cb5ac64f6db592cfbc8cc56f48a444378e2309 \ + --query "facebook" ``` ## Scope and trust -This first catalog is deliberately small. Its disclosed policy uses one automated -evidence-gate reviewer group rather than claiming human-review quorum. Fresh exact -endpoint observations are required, and source conflicts or dangerous drift need -two groups, so the single automated reviewer must abstain on those risks. Sticky -revocations and publisher/reviewer key separation remain enabled. +The v0.6.2 catalog contains 976 entities, 1,062 official-site edges, 20,178 +multilingual name facts, and Common Crawl Web Graph evidence for 840 domains that +already had imported identity assertions. Graph authority can prioritize review +and disambiguation, but cannot create an identity, official-site assertion, +review, vote, or redirect. The archive includes all 33 authenticated cold audit +objects referenced by the compact runtime generation. + +The bounded Wikidata discovery input is broad but not a representative or +high-demand sample. Its query and selection metadata are included for audit; raw +discovery output is never approval. + +The disclosed policy uses one automated evidence-gate reviewer group rather than +claiming human-review quorum. Fresh exact endpoint observations are required, and +source conflicts or dangerous drift need two groups, so the single automated +reviewer must abstain on those risks. Sticky revocations and +publisher/reviewer-key separation remain enabled. Consumers decide whether this policy is appropriate for their use. Preserve typed abstentions, retain attribution, and apply independent malware and content policy. @@ -60,3 +71,12 @@ The generation's approvals expire. Installing an immutable archive is not a prom that every decision stays valid forever: use the resolver's requested time, consume cumulative signed revocation feeds when published, and move to a newly signed full generation before relying on renewed decisions. + +## Regular updates + +The public source repository includes the same updater used to refresh candidate +generations. The example systemd timer runs weekly. It can download, authenticate, +import and build, but it holds no publisher key and cannot approve, sign or +activate a candidate. That separation lets any consumer automate evidence updates +without allowing a compromised downloader or changed upstream dataset to silently +change redirects.