ditto GitHub

ditto · did the data survive?

ditto compares two MongoDB collections and tells you how similar they are, with a GREEN YELLOW RED verdict and the metrics behind it. It works on raw BSON, so it knows nothing about your documents and works with any schema.

Built for one situation in particular: a batch job has just replaced a collection, and you have the previous state in a backup. Can you trust the new data, or should you roll back?

Maven Centralv0.5.0 Java 21+Spring Boot 4.1MongoDB 4.4+ Zero configuration to startLibrary + CLIStreaming, constant memoryLearns from historyMIT licensed

How it works

Documents are matched by key, compared in a canonical form, and profiled for structure, all in one streaming pass. Four kinds of signal feed the verdict.

01

Keys

Matched, added and removed documents. keySimilarity = matched / (matched + added + removed).

02

Content

SHA-256 of a canonical form: sorted fields, arrays as multisets, numbers by value. Field order and 42 vs 42.0 don't count as changes.

03

Changed paths

Which fields changed, how often, and from what to what, e.g. items[].price in 3% of documents, 19.9 → 0. One field changing everywhere means a broken mapping.

04

Structure

A path → BSON type histogram per side: vanished fields, new fields, and type shifts like int → double.

Each signal gets a level from configurable thresholds, and the overall verdict is the worst level. The report records which rule fired and why, with example keys you can paste straight into mongosh.

Zero configuration to start

ditto guesses what it can guess safely, and tells you about everything else.

ConventionDefault behaviourChange with
Backup namingcompareWithBackup("products"): products_backup → productsbackup-suffix
KeyDocuments matched by _idkey-field
Type hintsSpring Data's _class is ignoredalways-ignored-paths
ModeAUTO: full scan up to 5 million documents per side, a 20,000-key sample abovemode, full-scan-limit
MapsObjects with dynamic keys are detected and reported as attributes.*wildcard-paths
ThresholdsSensible defaults; learned from history once reports are storedthresholds.*
Everything elseReported as hints with ready-to-paste configuration–

ditto never ignores a field just because of its name, like updatedAt, because that could hide a sync that stopped working. It suggests, you decide. Every convention it applied is listed in the report under run.decisions.

Get started

Embed the library in a Spring Boot application, or run the CLI against any MongoDB.

  1. Add the dependency. It is on Maven Central. Your application needs Spring Boot 4 with MongoDB configured (spring.mongodb.*). Auto-configuration does the rest.

    pom.xml
    <dependency>
      <groupId>dev.jbaby.ditto</groupId>
      <artifactId>ditto-core</artifactId>
      <version>0.5.0</version>
    </dependency>

    Gradle: implementation("dev.jbaby.ditto:ditto-core:0.5.0").

  2. Inject CollectionComparator and compare. No configuration needed:

    Java
    @Service
    class ReplicationCheck {
    
        private final CollectionComparator comparator;
    
        ReplicationCheck(CollectionComparator comparator) {
            this.comparator = comparator;
        }
    
        Level verify() {
            ComparisonReport report = comparator.compareWithBackup("products");  // products_backup → products
            report.hints().forEach(hint -> log.info(hint.message()));          // what to configure next
            return report.verdict();                                            // GREEN, YELLOW or RED
        }
    }
  3. Follow the hints. The first report tells you what your data needs, e.g. a sync timestamp to ignore. Add it once:

    application.yml
    ditto:
      ignored-paths: [meta.syncedAt]
      expected-change-paths: [price, stock]

Options can also be set per request: comparator.compare(comparator.backupRequest("products").ignoredPaths("meta.syncedAt").build()). Every auto-configured bean is @ConditionalOnMissingBean, so you can replace any of them.

Usage scenarios

The same comparison serves several jobs. Open a scenario to see how to set it up.

Verify a batch replication, roll back on REDbackup vs. freshly synced collection

The batch job copies products to products_backup, then upserts products and deletes stale documents. Run the comparison as a separate step afterwards, and react to the verdict:

Java
ComparisonReport report = comparator.compare(
        ComparisonRequest.builder("products_backup", "products")
                .ignoredPaths("meta.syncedAt")
                .build(),
        progress -> log.info("{} docs read", progress.processedDocs()));

switch (report.verdict()) {
    case GREEN  -> log.info("products verified");
    case YELLOW -> alerts.warn("products", report.id());     // usable; have someone look
    case RED    -> rollback.restoreFromBackup("products");   // your code, e.g. renameCollection
}
  • A RED verdict is a value, not an exception. Exceptions mean the comparison couldn't run meaningfully: ComparisonException for a missing collection or unusable keys, DataAccessException for database errors.
  • Run it from any scheduler (JobRunr, Spring Batch, @Scheduled, Quartz). It's blocking, and stops when its thread is interrupted. The ProgressListener can drive a dashboard progress bar.
  • Label the comparison with the job's id, .label("batchJobId", jobId), and the stored report can be found from the job later: reportRepository.findByLabels(Map.of("batchJobId", jobId), 10) or GET /actuator/comparisons?label=batchJobId:4711.
  • Don't let the next sync overwrite the backup before verification and rollback are done.
Get a quick verdict on a huge collectionSAMPLE mode with confidence intervals

FULL mode reads both collections once. When that takes too long, sample instead: ditto draws random keys from both sides ($sample), looks up their counterparts in batches, and reports every rate as a 95% Wilson confidence interval.

ComparisonRequest.builder("events_backup", "events").sample(20_000).build();

# CLI
java -jar ditto-cli.jar ... --sample-size=20000

By default (verdict-basis: CONSERVATIVE) each rule evaluates the worse end of the interval, so GREEN means "GREEN with 95% confidence". A value right at a threshold may therefore come out YELLOW. Use a larger sample, or POINT, if you prefer point estimates.

A common pattern: a quick SAMPLE verdict right after the sync, then a FULL comparison later.

Check a test environment that holds only part of the datamatched-only

On a test environment one side often has just a few hundred documents, while the other is a full copy. A normal comparison is RED then, because most keys are missing. Matched-only compares just the documents whose key exists on both sides:

comparator.compare(comparator.backupRequest("products").matchedOnly().build());

# CLI
java -jar ditto-cli.jar --collection=products --matched-only
  • ditto reads the smaller collection and looks up its keys in the other, so a small test collection is quick to check against a large one.
  • Content, changed paths and structure cover the matched documents only. The key counts are still reported, but key similarity isn't judged, unless no key matches at all.
  • A normal comparison that looks like this situation suggests the option in its hints.
Validate a migration or a new pipelineold vs. new database, custom key

Compare the output of a rewritten pipeline with the old one, or a migrated database with its source. The two sides can live in different databases of the same cluster, and can be matched by any unique top-level field:

comparator.compare(ComparisonRequest.builder(
                CollectionRef.of("legacy", "customers"),
                CollectionRef.of("crm", "customers"))
        .keyField("customerNo")              // unique on both sides; index it
        .ignoredPaths("_id", "migratedAt")
        .nullEqualsMissing(true)             // new pipeline omits nulls
        .build());

Look at topChangedPaths for systematic differences: a path changed in close to 100% of documents points to a format or mapping difference rather than data churn.

Catch schema driftvanished fields and type shifts

Some changes keep every value intact but still break consumers. For example, an upstream system starts writing qty as a double. Content-wise 42 == 42.0, so the documents count as unchanged, but the structure check reports a type shift and the verdict is RED:

report.json
"typeShifts": [ {
  "path": "qty",
  "baseline":  [ { "type": "INT32",  "docs": 20000, "share": 1.0 } ],
  "candidate": [ { "type": "DOUBLE", "docs": 20000, "share": 1.0 } ]
} ]

The same check reports fields that disappeared from every document (missingPaths, RED), new fields (YELLOW) and fields that became much rarer or more common (presenceDeltas, YELLOW).

Gate a deployment or scriptCLI exit codes and jq

The CLI maps the verdict to the exit code and prints nothing but JSON on stdout, so it fits into shell scripts and CI pipelines:

bash
java -jar ditto-cli.jar --uri="$MONGO_URI" --db=shop \
  --baseline=products_backup --candidate=products --ignore=meta.syncedAt \
  --out=report.json > /dev/null
case $? in
  0) echo "GREEN" ;;
  1) echo "YELLOW, continuing"; jq '.rules[] | select(.level != "GREEN") | .reason' report.json ;;
  2) echo "RED, aborting"; exit 1 ;;
  *) echo "comparison failed"; exit 1 ;;
esac
Collections that change on purposeexpected-change paths and value examples

Prices, stock levels and counters change in most documents on every run. Left as they are, they would make the verdict RED. Declare them as expected-change paths. They stay in the report, flagged "expected": true, but don't count for maxPathChangeRate. Documents changed only there count as unchanged for the unchangedRate rule.

application.yml
ditto:
  expected-change-paths: [price, stock]    # stock also covers stock.qty, stock.updatedAt …
  redacted-paths: [customer]               # values shown as *** in reports
  max-value-examples: 3                    # 0 keeps values out of reports

Every changed path comes with a few before/after values, so a surprise is easy to judge:

report.json
{ "path": "name", "changedDocs": 1204, "expected": false,
  "valueExamples": [
    { "key": { "type": "INT32", "value": "4711" },
      "baseline": [ "\"Oak desk\"" ], "candidate": [ "\"OAK DESK\"" ] } ] }
Learn what's normal for each collectionthresholds from history

Fixed thresholds fit some collections badly: one where 8% of documents change every day is permanently YELLOW. Store every report, and let ditto derive the thresholds from the last runs of the same comparison:

application.yml
ditto:
  persistence:
    enabled: true                     # stores reports in comparison_reports
  adaptive-thresholds:
    enabled: true                     # after 5 runs: GREEN within 2σ, YELLOW within 3σ of the mean

ditto uses the last 20 non-RED runs, so a broken run doesn't lower the bar for the next one. The report says where its thresholds came from:

report.json
"thresholdSource": {
  "kind": "HISTORY", "historyRuns": 14,
  "notes": [ "unchangedRate: mean 0.9120, standard deviation 0.0071 over 14 runs -> GREEN >= 0.8978, YELLOW >= 0.8907", ... ]
}
Alert and monitorevents, Micrometer metrics, Actuator endpoint

Every comparison publishes a Spring event, whoever triggered it. Alerting is one listener:

Java
@EventListener
void onComparison(ComparisonCompletedEvent event) {
    if (event.report().verdict() == Level.RED) {
        alerts.critical(event.report());
    }
}
// ComparisonFailedEvent for comparisons that could not run

If the application has Micrometer, ditto records ditto.comparison (timer, tagged with verdict), ditto.comparison.errors, and gauges with the latest verdict, key similarity, unchanged rate and max path change rate per collection pair. That's enough for a dashboard and an alert rule.

With Actuator and persistence, the read-only comparisons endpoint lists stored reports (/actuator/comparisons, /actuator/comparisons/{id}) once exposed via management.endpoints.web.exposure.include. Both integrations are optional: ditto doesn't bring Micrometer or Actuator into an application that doesn't use them.

Configuration

You can start with none of it. The essentials below are what real data typically needs, and the report's hints tell you which ones. Set them as ditto.* properties, or per request; request options win.

Essentials

PropertyRequest optionDefaultMeaning
ignored-pathsignoredPaths–Technical fields removed before comparing, e.g. sync timestamps
expected-change-pathsexpectedChangePaths–Fields that are supposed to change (prices, counters): reported, not counted against the verdict
key-fieldkeyField_idTop-level field matching documents. Must be unique on both sides
modefullScan() / sample(n)AUTOAUTO, FULL or SAMPLE
persistence.enabled–falseStore every report; also switches on thresholds learned from history
thresholds.key-similarity.green / .yellowthresholds0.99 / 0.97GREEN if ≥ green, YELLOW if ≥ yellow, else RED
thresholds.unchanged-rate.green / .yellow0.95 / 0.85Same: higher is better
thresholds.max-path-change-rate.green / .yellow0.05 / 0.20GREEN if < green, YELLOW if < yellow, else RED

Path syntax

PatternMatches
meta.syncedAtA nested field; also everything below it
items[].priceField price of every element of array items; matrix[][] for nested arrays
attributes.** matches exactly one field name, e.g. attributes.color
Advanced optionsrarely needed; the defaults fit most data

Paths and keys

PropertyDefaultMeaning
wildcard-pathsdetectedMaps with dynamic keys, ….*; usually detected automatically
order-sensitive-paths–Arrays whose element order matters
redacted-paths–Values shown as *** in value examples
always-ignored-paths_classIgnored in every comparison, on top of ignored-paths
null-equals-missingfalseTreat field: null like an absent field
matched-onlyfalseCompare only documents whose key exists on both sides, e.g. on a test environment
mixed-key-typesREJECTCOMPARE accepts keys of different BSON types
backup-suffix_backupBaseline name for compareWithBackup / --collection

Mode and detection

PropertyDefaultMeaning
full-scan-limit5000000AUTO scans fully up to this many documents per side
sample.size20000Keys sampled per side
sample.verdict-basisCONSERVATIVEEvaluate the worse confidence bound, or the POINT estimate
sample.lookup-batch-size500Keys per $in lookup
map-detection.enabledtrueDetect maps with dynamic keys before comparing
map-detection.sample-size / .min-distinct-keys500 / 20Documents sampled per side / distinct names needed for a map

Thresholds and history

PropertyDefaultMeaning
thresholds.structure.type-share-delta0.01Type-share change that counts as a type shift (RED)
thresholds.structure.presence-delta0.05Presence changes beyond this are YELLOW
thresholds.structure.vanished-min-presence0.0Vanished paths rarer than this are YELLOW instead of RED
adaptive-thresholds.enabled= persistenceDerive thresholds from stored reports
adaptive-thresholds.history-size / .min-history20 / 5Previous non-RED runs considered / needed
adaptive-thresholds.green-sigma / .yellow-sigma2.0 / 3.0Band distance from the historical mean, in standard deviations
adaptive-thresholds.min-spread0.005Lower limit for the standard deviation

Report, persistence and resources

PropertyDefaultMeaning
persistence.collection / .databasecomparison_reports / defaultWhere reports are stored
persistence.retentionforeverHow long reports are kept, e.g. 365d (TTL index)
labels.<key>–Labels stored with every report, e.g. labels.environment=test
max-examples20Example keys per category and per changed path
max-value-examples3Before/after values per changed path; 0 disables them
top-changed-paths50Changed paths listed in the report
max-tracked-paths10000Distinct paths tracked per side; bounds memory
batch-size1000Cursor batch size
no-cursor-timeoutfalseKeep idle server cursors alive on very large collections
progress-interval10sProgress logging and listener interval

The connection itself uses Spring Boot's spring.mongodb.* properties.

Reading a report

Start at the top and work down: verdict, the rules that fired, then the metrics that explain them.

In the browser

--html=report.html writes the report as a single HTML page next to the JSON: verdict, hints with ready-to-copy options, rules, changed paths with before/after values and the structure diff. The page makes no external requests, so it works offline, as a CI artifact or as an e-mail attachment. For a stored report, use java -jar ditto-cli.jar report --label=batchJobId=4711 --html=report.html; in your own code, inject ReportHtml.

Already have a report.json? Open it in the report viewer: it renders the file in your browser and uploads nothing. See an example report.

The JSON

report.json (abbreviated)
{
  "verdict": "YELLOW",
  "rules": [
    { "rule": "keySimilarity", "level": "GREEN", "observed": 0.9950,
      "reason": "0.9950 meets GREEN (>= 0.9900) (matched 19900, added 50, removed 50)" },
    { "rule": "maxPathChangeRate", "level": "YELLOW", "observed": 0.0812,
      "reason": "Path 'price' changed in 0.0812 of matched documents, at or above the GREEN limit 0.0500 ...",
      "details": [ "price: 0.0812 (1616 documents)" ] },
    ...
  ],
  "keys":            { "matched": 19900, "added": 50, "removed": 50, "keySimilarity": { "kind": "exact", "value": 0.995 } },
  "content":         { "unchanged": 18284, "changed": 1616, "unchangedRate": { "value": 0.9188 } },
  "topChangedPaths": [ { "path": "price", "changedDocs": 1616, "expected": false,
                         "examples": [ { "type": "OBJECT_ID", "value": "{\"$oid\": \"6553f1212161972337cc2db4\"}" } ],
                         "valueExamples": [ { "baseline": [ "412.07" ], "candidate": [ "0" ], ... } ] } ],
  "structure":       { "newPaths": [], "missingPaths": [], "typeShifts": [], "presenceDeltas": [] },
  "examples":        { "changed": [...], "added": [...], "removed": [...] },
  "run":             { "durationMillis": 686, "baselineCount": 19950, "settings": { ... } },
  "warnings":        [],
  "hints":           [ { "kind": "EXPECTED_CHANGE", "path": "price",
                         "message": "price changed in 8% of matched documents. If this field is supposed to change …",
                         "property": "ditto.expected-change-paths=price", "cliOption": "--expected=price" } ]
}
If you see…It usually means…
Many removed, few addedA partial or aborted load
Many removed and many addedKey values changed format ("123" vs 123) or a new ID scheme
One path changed in ~100% of documentsA mapping/format change in that field, or a technical field that should be ignored
Many paths at low ratesNormal data churn
Fields that change in most documents every runLegitimate churn: declare them as expected-change-paths
Anything in hintsA ready-to-paste property or CLI option for the next run
A type shift with unchanged contentSame values, different BSON type: consumers that read typed values may break
warnings about the path capA map with dynamic keys: add it to wildcard-paths

Example keys are relaxed Extended JSON, ready for mongosh: db.products.find({_id: {"$oid": "6553f1212161972337cc2db4"}}). Look the same key up in both collections to see the difference.

CLI reference

Short options map onto the request. Any --ditto.* or --spring.mongodb.* property also works.

OptionMeaning
--uri, --dbConnection string (default mongodb://localhost:27017) and default database (default test)
--collectionCompares <collection>_backup with <collection>
--baseline, --candidateOr name both collections explicitly
--baseline-db, --candidate-dbDatabase per side, if not --db
--keyKey field
--ignore, --ordered, --wildcard, --expected, --redactComma-separated path lists
--mode=auto|full|sample, --sample-size=nComparison mode (default auto)
--matched-onlyCompare only documents whose key exists on both sides
--null-equals-missing, --mixed-key-types, --verdict-basisAs in the configuration table
--label=key=valueLabel stored with the report, e.g. a batch job id; repeatable
--persistStore the report in MongoDB
--out=fileAlso write the report to a file
--html=fileAlso write the report as a self-contained HTML page
report --id=…, --label=…, --collection=…Print a stored report (and with --html write it as a page); the newest one if several match. Exits with its verdict
--helpFull usage

Test-data generator

java -jar ditto-cli.jar generate writes a baseline and a modified copy of product-like documents. They include nested objects, arrays of objects, a dynamic-key map, an ordered event list, decimals, dates and nullable fields.

OptionDefaultMeaning
--baseline / --candidatedemo_backup / demoCollection names (dropped and recreated)
--docs10000Baseline size
--seed42Same seed, same data
--touch-synctrueRefresh meta.syncedAt in every candidate document, like a real sync
--changes–Comma-separated kind[:path]:fraction, e.g. modify:price:0.05,drop-field:legacyCode:1,int-to-double:qty:0.5,shuffle-arrays:1,delete-docs:0.01,add-docs:0.01

Good to know

More detail in the README, including limitations and a sketch of a JobRunr integration with rollback.