data2dsl

data2dsl is a planned, evidence-first comparison layer. It will turn facts from existing data sources into comparable observations and deterministic differences that other systems can reason about.

The short version:

Ask one bounded question, acquire the relevant facts from two or more sources, normalize them without losing provenance, compare like with like, and return the result together with evidence.

The project is currently in contract and integration planning. The repository contains governance, capability evidence and architectural decisions, but no functional product implementation or final public DSL yet.

The problem

Useful facts already exist across Markdown documents, Git repositories, GitHub, configuration files, code analyzers and browser-backed sources. Each source has its own structure and vocabulary. Today a consumer such as todo2code must either understand every source or rely on an LLM to interpret incomparable outputs.

That creates four recurring problems:

  1. the same metric can be named or represented differently by each source;
  2. values may refer to different actors, repositories or time windows;
  3. conclusions can lose the evidence needed to verify them;
  4. source acquisition, deterministic comparison and higher-level reasoning get mixed into one component.

data2dsl is intended to provide the missing factual boundary between source tools and reasoning consumers.

Who it is for

The primary consumers are programs and agents that need to compare claims with observed data while preserving provenance. Initial consumers are expected to include todo2code and repository-governance workflows, but the core must not depend on either one.

A human may formulate the question, inspect the differences and follow the evidence. A source adapter acquires facts. data2dsl normalizes and compares them. A separate consumer decides what the result means or what action, if any, should follow.

Golden case

The first end-to-end case is:

Compare statements in work-summary.md with actual GitHub activity for the same repository, actor, metric and time window.

For example, a summary might claim 12 commits for a person during a given week, while the GitHub source reports 10. The planned result is not prose or an LLM verdict. It is an evidence-bearing comparison containing, conceptually:

Field Example
Subject repository and actor
Metric commit count
Window explicit start and end
Left observation claimed value from a Markdown location
Right observation measured value from GitHub pages/API results
Outcome CONFLICT
Delta -2
Evidence immutable references and content digests for both sides

This table illustrates intended behavior; it is not a final API or schema.

Planned inputs

A bounded comparison needs three kinds of input:

Natural-language interpretation may help construct a query, but it must not silently change the metric, window or source identity. Unresolved ambiguity must remain visible.

Planned outputs

The factual output should contain:

UNEVALUABLE is not success and missing data is not zero. Comparison outcomes are also distinct from the state of an individual observation.

Planned composition

flowchart LR
    Q["Bounded query"] --> R["Routing and explicit mapping"]
    R --> M["Markdown via mdflow"]
    R --> G["Git factual seam"]
    R --> H["GitHub via Diagit extension"]
    R --> C["Existing code/data analyzers"]
    M --> O["Comparable observations + evidence"]
    G --> O
    H --> O
    C --> O
    O --> D["Deterministic comparator"]
    D --> F["Facts, outcomes, deltas, gaps, evidence"]
    F --> X["todo2code or another reasoning consumer"]

This is a composition hypothesis, not a final runtime contract. Current reuse decisions and their pinned evidence are recorded in docs/CAPABILITY_MAP.md.

What data2dsl owns

The project should own only the smallest missing responsibilities:

What data2dsl does not own

The project is not intended to become:

Source adapters remain responsible for truthful acquisition. Standards owners remain responsible for shared contracts. Consumers remain responsible for reasoning, policy and action.

Reuse-first strategy

Every capability follows this order:

  1. REUSE an existing public API or CLI when its behavior and ownership fit.
  2. EXTRACT the smallest neutral seam when useful behavior is trapped inside another product; preserve its language and compatibility.
  3. EXTEND the established owning component when a nearby capability exists.
  4. Mark a capability MISSING and implement it locally only after the first three options have been disproved with current evidence.

Examples from the Phase 0 inventory include reusing mdflow for Markdown structure, extending Diagit’s established GitHub boundary for commit metrics, and keeping todo2code as a reasoning consumer rather than moving its policy into data2dsl.

Delivery roadmap

The planned delivery order is dependency-driven:

  1. decide the observation/evidence contract and its compatibility with subactor/twin;
  2. agree a minimal shared query/result profile with its standards owner;
  3. define the smallest deterministic scalar/set comparison semantics;
  4. extend Diagit with the read-only GitHub metrics required by the golden case;
  5. implement and validate work-summary.md versus GitHub in Docker;
  6. evaluate Git/config/AST extraction from todo2code only when a real second consumer proves it is necessary;
  7. integrate factual results back into todo2code without moving reasoning into data2dsl.

Each step requires its own bounded ticket and evidence. Changes to another repository require that repository’s owner-approved workflow.

Current state

See TODO.md for current work and project/TICKETS.md for governed evidence.

Governance

This repository adopts an immutable published revision of wellmanifest/new-project. Multi-step work is ticket-governed and bounded by the active ticket’s intent.json. Human-owned user-* files are never written by agents, and implementation claims require deterministic validation rather than README text alone.