Files
documents/dsearch.md
T
2026-09-17 08:40:45 +00:00

13 KiB

DSEARCH

Purpose

DSEARCH is the architectural and specification effort directed toward making a digital search commons possible. It provides shared search-domain semantics and interoperability boundaries through which independently built and independently operated systems can exchange useful discovery and search information while retaining their own infrastructure, internal terminology, methods, policies, and implementations.

The commons is larger than DSEARCH, and DSEARCH is larger than any one implementation. Crawlers, archives, relays, databases, protocols, applications, AI systems, specialist indexes, search engines, researchers, and infrastructure operators can each supply useful capabilities without any one of those components becoming the definition of the commons.

DSEARCH is concerned especially with the relationships that allow those capabilities to become useful across system boundaries.

Domain-First Architecture

DSEARCH begins with agreements about meaning. Those agreements define such questions as what constitutes a resource, what an observation represents, what replication accomplishes, what coverage means, how currentness differs from timestamp precision, and what provenance must remain visible for search use.

Implementation technologies then carry those meanings.

commons requirements
        ↓
Base Layer contracts
        ↓
domain/profile contracts
        ↓
bindings and adapters
        ↓
implementations

The intended relationship is that the commons defines meaning while implementation technologies supply capability where doing so reduces total complexity.

A substrate may supply valuable machinery for identity, serialization, transport, storage, synchronization, search, discovery, or deployment. DSEARCH should use such capabilities without allowing substrate-native identifiers, records, tags, schemas, or operating assumptions to silently redefine the search domain.

This produces a standing replaceability test: a foundational concept should remain describable independently of the technology currently carrying it.

Participants and Specialization

DSEARCH assumes heterogeneous participation. One participant may crawl the broad web. Another may maintain a specialist index. Another may preserve historical representations. Another may contribute ranking, storage, entity search, citation analysis, relay capacity, retrieval interfaces, or AI-assisted analysis.

Participants need not implement every role. A useful contribution becomes part of the commons when its meaning can be understood and used outside the implementation that produced it.

The desired condition is diversity with interoperability.

DSEARCH and DEVIDENCE

DSEARCH and DEVIDENCE share concepts around resources, observers, observations, provenance, and evidence but emphasize different responsibilities.

DSEARCH is optimized around questions such as:

What resource is this?
How can it be found?
Who observed or described it?
Which indexes know about it?
Which paths can retrieve the observation?
How current is the available evidence?
What is the coverage of this participant?
Where has the observation been replicated?
Which search route produced this result?

DEVIDENCE is optimized around deeper questions such as what exact representation was observed, how it was acquired, how an analytical observation was derived from it, what historical versions exist, and why two observers or extractors disagree.

DSEARCH must be able to consume both thin and rich contributions. A signed SIP-01 event can be useful search evidence even when no richer DEVIDENCE record exists. When richer evidence is available, DSEARCH should be able to expose or reference it without requiring every participant to implement the same archival system.

Search Results and Evidence Depth

A DSEARCH result can point to a resource while retaining enough context to identify the observation that caused the resource to be searchable.

Where richer evidence is available, the result may also expose relationships to preserved representations, acquisition records, derivations, corroborating or conflicting observations, historical versions, and replication state.

This should be treated as evidence depth, not as a universal trust score.

thin result
  resource + observer + searchable observation

richer result
  + preserved representation
  + acquisition provenance
  + derivation lineage

richer still
  + independent observations
  + historical versions
  + agreement / disagreement evidence

The consumer remains responsible for evaluation according to its own purpose.

Coexisting Search Surfaces

DSEARCH explicitly permits multiple searchable projections to coexist.

A lightweight SIP-compatible surface may publish web observations as kind 39697 to ordinary Nostr relays and permit retrieval through NIP-50-compatible clients. A richer DSEARCH/DEVIDENCE surface may index resources, representations, entities, citations, claims, historical versions, and evidentiary relationships.

Neither surface is a migration stage toward the other.

                       shared evidence/resource base
                                  |
                    +-------------+-------------+
                    |                           |
                    v                           v
             SIP-01 projection          DSEARCH projection
             lightweight                richer resource and
             interoperable              evidence search
                    |                           |
                    v                           v
              Nostr relays              DSEARCH indexes/APIs

A third-party SIP indexer that knows nothing about DEVIDENCE can remain a valid participant. A richer participant can publish the same ordinary SIP projection while separately preserving the deeper evidence chain.

DSEARCH is not limited to current live-web observations. Historical representations imported from WARC, ARC, WACZ, archival services, local collections, or other sources can become searchable resources or support searchable analytical observations.

Historical evidence must not be represented as current evidence merely because it was ingested recently.

DSEARCH therefore relies on the Terms of Reference distinction between observed time and ingested time, and between a historical representation and a current resource observation.

A preserved archived representation can be searchable in its own right even when the original live resource is gone.

Currentness and Coverage

Currentness is an evidentiary conclusion rather than a timestamp alias. A precisely dated observation establishes that something was observed at a particular time. Later confirming observations may strengthen the case that it still describes the resource. Historical observations can remain valuable without being presented as evidence of current state.

Coverage likewise needs a bounded meaning. A participant may describe or demonstrate the resource class, geography, language, time period, topic, shard range, corpus, or other domain in which its observations or search capabilities apply.

Coverage information can also support decentralized crawl coordination. If several observers can determine that a lightly resourced site was recently and adequately observed, they can reduce redundant acquisition pressure without requiring a central scheduler.

Publication, Replication, and Independent Observation

DSEARCH preserves the distinction between publication, replication, and independent observation.

Publication makes a record available. Replication creates additional copies or retrievable paths. Independent observation is a separate evidentiary act.

The distinction protects the commons against confused evidence. A widely replicated event may be highly available without becoming independently corroborated. Many valid signatures may represent several keys without proving several independent operators. Search interfaces should retain enough provenance to allow these differences to remain visible.

Decentralization

DSEARCH treats decentralization as instrumental to the health of the commons rather than as an undifferentiated numerical objective.

Infrastructure decentralization concerns independent failure domains and alternative operational paths for crawling, storage, indexing, communication, search, and preservation.

Implementation decentralization concerns the ability of independently developed systems to satisfy the same contracts without inheriting the internal architecture of a reference implementation.

Governance decentralization concerns the balance between the specialized stewardship required to maintain coherent contracts and the operational independence retained by participants.

Economic decentralization concerns preserving viable participation without making one commercial position, paid service, token system, or provider a compulsory gate.

Evidentiary decentralization concerns plurality of observation rather than plurality of copies. Many servers or keys do not establish plural evidence if they derive their information from one underlying source or administrative authority.

Concentration is not automatically a defect. A successful implementation, transport, or service can legitimately become widely used. The architectural boundary is capture: success in one layer should not acquire compulsory authority over meanings or participation that belong to the wider commons.

Resilience and Coexistence

DSEARCH distinguishes technical, architectural, and governance resilience.

Technical resilience allows useful operation to continue through local failures.

Architectural resilience allows implementation mechanisms to be replaced while preserving shared meanings.

Governance resilience allows contracts and specifications to evolve while independently controlled participants migrate on different schedules.

These forms of resilience require coexistence. Several versions, implementations, policies, transports, and operating approaches may remain active during the same period.

Deprecation expresses architectural direction; migration changes operating implementations. They are therefore separate processes.

Threats and Defensive Principles

A digital search commons faces threats from both malicious behavior and ordinary success.

Manufactured plurality can use many identities to simulate independence. Signatures should therefore establish provenance without being confused with operator independence or truth.

Substrate capture can occur when implementation-specific records, tags, identifiers, or query assumptions gradually become the only practical language in which the commons can express its own concepts.

Economic capture can occur when a successful service becomes such a compulsory dependency that its business rules begin to determine the architecture.

Fragmented security boundaries can allow external input, shared credentials, mutation authority, or assumptions about another participant's availability to propagate failures across otherwise independent systems.

Resource exhaustion can target the commons directly through traffic and data volume, while the commons itself can inadvertently become an attacker when many independent crawlers repeatedly observe the same small resource.

Assumed erasure can confuse withdrawal from one participant with disappearance from the wider commons.

The recurring defenses are explicit semantics, bounded authority, provenance, containment, replaceable adapters, conformance tests, versioning, migration paths, observation etiquette, and preservation of independence.

Empirical Accretion

DSEARCH should refine its Base Layer against real operating evidence. New concepts should not be promoted merely because one implementation exposes a convenient field.

The development method is:

real search/observation case
        ↓
represent with existing vocabulary
        ↓
identify material information loss
        ↓
collect additional examples
        ↓
introduce the smallest useful distinction
        ↓
test against other domains/implementations
        ↓
promote only if it generalizes

Housatonic web indexing is the first substantial corpus intended to exercise this method. It should pressure the model around aliases, historical versions, redirects, currentness, archival backfill, independently derived metadata, relay replication, and richer evidence search without making Housatonic-specific structures part of the Base Layer.

Layered Specification Work

The continuing DSEARCH work should remain separated into distinct levels of authority:

vision and commons principles
        ↓
Terms of Reference
        ↓
Base Layer invariants
        ↓
contracts
        ↓
schemas
        ↓
conformance tests and vectors
        ↓
bindings and adapters
        ↓
reference implementations
        ↓
applications and participants

Conformance should test observable meaning rather than incidental source structure. Independent implementations provide evidence that a contract has meaning outside the codebase that first expressed it.