ditto · did the data survive?
ditto compares two MongoDB collections and tells you how similar they are, with a GREEN YELLOW RED verdict and the metrics behind it. It works on raw BSON, so it knows nothing about your documents and works with any schema.
Built for one situation in particular: a batch job has just replaced a collection, and you have the previous state in a backup. Can you trust the new data, or should you roll back?
How it works
Documents are matched by key, compared in a canonical form, and profiled for structure, all in one streaming pass. Four kinds of signal feed the verdict.
Keys
Matched, added and removed documents. keySimilarity = matched / (matched + added + removed).
Content
SHA-256 of a canonical form: sorted fields, arrays as multisets, numbers by value. Field order and
42 vs 42.0 don't count as changes.
Changed paths
Which fields changed, how often, and from what to what, e.g. items[].price in 3% of
documents, 19.9 → 0. One field changing everywhere means a broken mapping.
Structure
A path → BSON type histogram per side: vanished fields, new fields, and type shifts like
int → double.
Each signal gets a level from configurable thresholds, and the overall verdict is the worst level. The report records which rule fired and why, with example keys you can paste straight into mongosh.
Zero configuration to start
ditto guesses what it can guess safely, and tells you about everything else.
| Convention | Default behaviour | Change with |
|---|---|---|
| Backup naming | compareWithBackup("products"): products_backup → products | backup-suffix |
| Key | Documents matched by _id | key-field |
| Type hints | Spring Data's _class is ignored | always-ignored-paths |
| Mode | AUTO: full scan up to 5 million documents per side, a 20,000-key sample above | mode, full-scan-limit |
| Maps | Objects with dynamic keys are detected and reported as attributes.* | wildcard-paths |
| Thresholds | Sensible defaults; learned from history once reports are stored | thresholds.* |
| Everything else | Reported as hints with ready-to-paste configuration | – |
ditto never ignores a field just because of its name, like updatedAt, because that could hide a
sync that stopped working. It suggests, you decide. Every convention it applied is listed in the report under
run.decisions.
Get started
Embed the library in a Spring Boot application, or run the CLI against any MongoDB.
-
Add the dependency. It is on Maven Central. Your application needs Spring Boot 4 with MongoDB configured (
spring.mongodb.*). Auto-configuration does the rest.pom.xml<dependency> <groupId>dev.jbaby.ditto</groupId> <artifactId>ditto-core</artifactId> <version>0.5.0</version> </dependency>
Gradle:
implementation("dev.jbaby.ditto:ditto-core:0.5.0"). -
Inject
CollectionComparatorand compare. No configuration needed:Java@Service class ReplicationCheck { private final CollectionComparator comparator; ReplicationCheck(CollectionComparator comparator) { this.comparator = comparator; } Level verify() { ComparisonReport report = comparator.compareWithBackup("products"); // products_backup → products report.hints().forEach(hint -> log.info(hint.message())); // what to configure next return report.verdict(); // GREEN, YELLOW or RED } } -
Follow the hints. The first report tells you what your data needs, e.g. a sync timestamp to ignore. Add it once:
application.ymlditto: ignored-paths: [meta.syncedAt] expected-change-paths: [price, stock]
Options can also be set per request:
comparator.compare(comparator.backupRequest("products").ignoredPaths("meta.syncedAt").build()).
Every auto-configured bean is @ConditionalOnMissingBean, so you can replace any of them.
-
Download the jar from the latest release, or build it:
git clone https://github.com/danielbartl/ditto.git cd ditto && ./mvnw package -DskipTests # → ditto-cli/target/ditto-cli.jar -
Compare two collections. The report goes to stdout as JSON; logs and a one-line summary go to stderr.
# products_backup (baseline) vs. products (candidate) java -jar ditto-cli.jar --uri="mongodb://user:pass@host:27017" --db=shop \ --collection=products > report.jsonstderr shows the verdict, and hints with ready-made options for the next run:
Verdict RED: keySimilarity 0.9901, unchangedRate 0.0000, 2 changed paths, 1598 ms Hints: - meta.syncedAt changed in 100% of matched documents and holds dates: it looks like a technical timestamp written by every run. If so, ignore it. --ignore=meta.syncedAt -
Use the exit code.
Exit code Meaning 0GREEN 1YELLOW 2RED 3Error: bad options, missing collection, unusable keys, connection failure
Try ditto on generated data before pointing it at real collections. The repository includes a docker-compose MongoDB and a generator that writes a baseline plus a copy with the changes you choose.
docker compose up -d mongo ./mvnw package -DskipTests # 50k product-like documents; 3% prices changed, 0.5% deleted, 0.5% added SEED_ARGS="--docs=50000 --changes=modify:price:0.03,delete-docs:0.005,add-docs:0.005" \ docker compose run --rm seed java -jar ditto-cli/target/ditto-cli.jar --db=ditto --collection=demo # → RED, with a hint to ignore meta.syncedAt; then: java -jar ditto-cli/target/ditto-cli.jar --db=ditto --collection=demo --ignore=meta.syncedAt
Available changes: modify, drop-field, add-field, set-null,
int-to-double, to-string, shuffle-arrays, delete-docs,
add-docs. See the generator reference.
Usage scenarios
The same comparison serves several jobs. Open a scenario to see how to set it up.
Verify a batch replication, roll back on REDbackup vs. freshly synced collection
The batch job copies products to products_backup, then upserts
products and deletes stale documents. Run the comparison as a separate step afterwards, and
react to the verdict:
ComparisonReport report = comparator.compare(
ComparisonRequest.builder("products_backup", "products")
.ignoredPaths("meta.syncedAt")
.build(),
progress -> log.info("{} docs read", progress.processedDocs()));
switch (report.verdict()) {
case GREEN -> log.info("products verified");
case YELLOW -> alerts.warn("products", report.id()); // usable; have someone look
case RED -> rollback.restoreFromBackup("products"); // your code, e.g. renameCollection
}- A RED verdict is a value, not an exception. Exceptions mean the comparison couldn't run
meaningfully:
ComparisonExceptionfor a missing collection or unusable keys,DataAccessExceptionfor database errors. - Run it from any scheduler (JobRunr, Spring Batch,
@Scheduled, Quartz). It's blocking, and stops when its thread is interrupted. TheProgressListenercan drive a dashboard progress bar. - Label the comparison with the job's id,
.label("batchJobId", jobId), and the stored report can be found from the job later:reportRepository.findByLabels(Map.of("batchJobId", jobId), 10)orGET /actuator/comparisons?label=batchJobId:4711. - Don't let the next sync overwrite the backup before verification and rollback are done.
Get a quick verdict on a huge collectionSAMPLE mode with confidence intervals
FULL mode reads both collections once. When that takes too long, sample instead: ditto draws random keys
from both sides ($sample), looks up their counterparts in batches, and reports every rate as a
95% Wilson confidence interval.
ComparisonRequest.builder("events_backup", "events").sample(20_000).build();
# CLI
java -jar ditto-cli.jar ... --sample-size=20000By default (verdict-basis: CONSERVATIVE) each rule evaluates the worse end of the
interval, so GREEN means "GREEN with 95% confidence". A value right at a
threshold may therefore come out YELLOW. Use a larger sample, or
POINT, if you prefer point estimates.
A common pattern: a quick SAMPLE verdict right after the sync, then a FULL comparison later.
Check a test environment that holds only part of the datamatched-only
On a test environment one side often has just a few hundred documents, while the other is a full copy. A normal comparison is RED then, because most keys are missing. Matched-only compares just the documents whose key exists on both sides:
comparator.compare(comparator.backupRequest("products").matchedOnly().build());
# CLI
java -jar ditto-cli.jar --collection=products --matched-only- ditto reads the smaller collection and looks up its keys in the other, so a small test collection is quick to check against a large one.
- Content, changed paths and structure cover the matched documents only. The key counts are still reported, but key similarity isn't judged, unless no key matches at all.
- A normal comparison that looks like this situation suggests the option in its hints.
Validate a migration or a new pipelineold vs. new database, custom key
Compare the output of a rewritten pipeline with the old one, or a migrated database with its source. The two sides can live in different databases of the same cluster, and can be matched by any unique top-level field:
comparator.compare(ComparisonRequest.builder(
CollectionRef.of("legacy", "customers"),
CollectionRef.of("crm", "customers"))
.keyField("customerNo") // unique on both sides; index it
.ignoredPaths("_id", "migratedAt")
.nullEqualsMissing(true) // new pipeline omits nulls
.build());Look at topChangedPaths for systematic differences: a path changed in close to 100% of
documents points to a format or mapping difference rather than data churn.
Catch schema driftvanished fields and type shifts
Some changes keep every value intact but still break consumers. For example, an upstream system starts
writing qty as a double. Content-wise 42 == 42.0, so the documents count as
unchanged, but the structure check reports a type shift and the verdict is
RED:
"typeShifts": [ {
"path": "qty",
"baseline": [ { "type": "INT32", "docs": 20000, "share": 1.0 } ],
"candidate": [ { "type": "DOUBLE", "docs": 20000, "share": 1.0 } ]
} ]The same check reports fields that disappeared from every document (missingPaths, RED),
new fields (YELLOW) and fields that became much rarer or more common (presenceDeltas, YELLOW).
Gate a deployment or scriptCLI exit codes and jq
The CLI maps the verdict to the exit code and prints nothing but JSON on stdout, so it fits into shell scripts and CI pipelines:
java -jar ditto-cli.jar --uri="$MONGO_URI" --db=shop \ --baseline=products_backup --candidate=products --ignore=meta.syncedAt \ --out=report.json > /dev/null case $? in 0) echo "GREEN" ;; 1) echo "YELLOW, continuing"; jq '.rules[] | select(.level != "GREEN") | .reason' report.json ;; 2) echo "RED, aborting"; exit 1 ;; *) echo "comparison failed"; exit 1 ;; esac
Collections that change on purposeexpected-change paths and value examples
Prices, stock levels and counters change in most documents on every run. Left as they are, they would make
the verdict RED. Declare them as expected-change paths. They
stay in the report, flagged "expected": true, but don't count for
maxPathChangeRate. Documents changed only there count as unchanged for the
unchangedRate rule.
ditto: expected-change-paths: [price, stock] # stock also covers stock.qty, stock.updatedAt … redacted-paths: [customer] # values shown as *** in reports max-value-examples: 3 # 0 keeps values out of reports
Every changed path comes with a few before/after values, so a surprise is easy to judge:
{ "path": "name", "changedDocs": 1204, "expected": false,
"valueExamples": [
{ "key": { "type": "INT32", "value": "4711" },
"baseline": [ "\"Oak desk\"" ], "candidate": [ "\"OAK DESK\"" ] } ] }Learn what's normal for each collectionthresholds from history
Fixed thresholds fit some collections badly: one where 8% of documents change every day is permanently YELLOW. Store every report, and let ditto derive the thresholds from the last runs of the same comparison:
ditto:
persistence:
enabled: true # stores reports in comparison_reports
adaptive-thresholds:
enabled: true # after 5 runs: GREEN within 2σ, YELLOW within 3σ of the meanditto uses the last 20 non-RED runs, so a broken run doesn't lower the bar for the next one. The report says where its thresholds came from:
"thresholdSource": {
"kind": "HISTORY", "historyRuns": 14,
"notes": [ "unchangedRate: mean 0.9120, standard deviation 0.0071 over 14 runs -> GREEN >= 0.8978, YELLOW >= 0.8907", ... ]
}Alert and monitorevents, Micrometer metrics, Actuator endpoint
Every comparison publishes a Spring event, whoever triggered it. Alerting is one listener:
@EventListener
void onComparison(ComparisonCompletedEvent event) {
if (event.report().verdict() == Level.RED) {
alerts.critical(event.report());
}
}
// ComparisonFailedEvent for comparisons that could not runIf the application has Micrometer, ditto records ditto.comparison (timer, tagged with
verdict), ditto.comparison.errors, and gauges with the latest verdict, key similarity,
unchanged rate and max path change rate per collection pair. That's enough for a dashboard and an alert
rule.
With Actuator and persistence, the read-only comparisons endpoint lists stored reports
(/actuator/comparisons, /actuator/comparisons/{id}) once exposed via
management.endpoints.web.exposure.include. Both integrations are optional: ditto doesn't
bring Micrometer or Actuator into an application that doesn't use them.
Configuration
You can start with none of it. The essentials below are what real data typically needs, and the report's hints
tell you which ones. Set them as ditto.* properties, or per request; request options win.
Essentials
| Property | Request option | Default | Meaning |
|---|---|---|---|
ignored-paths | ignoredPaths | – | Technical fields removed before comparing, e.g. sync timestamps |
expected-change-paths | expectedChangePaths | – | Fields that are supposed to change (prices, counters): reported, not counted against the verdict |
key-field | keyField | _id | Top-level field matching documents. Must be unique on both sides |
mode | fullScan() / sample(n) | AUTO | AUTO, FULL or SAMPLE |
persistence.enabled | – | false | Store every report; also switches on thresholds learned from history |
thresholds.key-similarity.green / .yellow | thresholds | 0.99 / 0.97 | GREEN if ≥ green, YELLOW if ≥ yellow, else RED |
thresholds.unchanged-rate.green / .yellow | 0.95 / 0.85 | Same: higher is better | |
thresholds.max-path-change-rate.green / .yellow | 0.05 / 0.20 | GREEN if < green, YELLOW if < yellow, else RED |
Path syntax
| Pattern | Matches |
|---|---|
meta.syncedAt | A nested field; also everything below it |
items[].price | Field price of every element of array items; matrix[][] for nested arrays |
attributes.* | * matches exactly one field name, e.g. attributes.color |
Advanced optionsrarely needed; the defaults fit most data
Paths and keys
| Property | Default | Meaning |
|---|---|---|
wildcard-paths | detected | Maps with dynamic keys, ….*; usually detected automatically |
order-sensitive-paths | – | Arrays whose element order matters |
redacted-paths | – | Values shown as *** in value examples |
always-ignored-paths | _class | Ignored in every comparison, on top of ignored-paths |
null-equals-missing | false | Treat field: null like an absent field |
matched-only | false | Compare only documents whose key exists on both sides, e.g. on a test environment |
mixed-key-types | REJECT | COMPARE accepts keys of different BSON types |
backup-suffix | _backup | Baseline name for compareWithBackup / --collection |
Mode and detection
| Property | Default | Meaning |
|---|---|---|
full-scan-limit | 5000000 | AUTO scans fully up to this many documents per side |
sample.size | 20000 | Keys sampled per side |
sample.verdict-basis | CONSERVATIVE | Evaluate the worse confidence bound, or the POINT estimate |
sample.lookup-batch-size | 500 | Keys per $in lookup |
map-detection.enabled | true | Detect maps with dynamic keys before comparing |
map-detection.sample-size / .min-distinct-keys | 500 / 20 | Documents sampled per side / distinct names needed for a map |
Thresholds and history
| Property | Default | Meaning |
|---|---|---|
thresholds.structure.type-share-delta | 0.01 | Type-share change that counts as a type shift (RED) |
thresholds.structure.presence-delta | 0.05 | Presence changes beyond this are YELLOW |
thresholds.structure.vanished-min-presence | 0.0 | Vanished paths rarer than this are YELLOW instead of RED |
adaptive-thresholds.enabled | = persistence | Derive thresholds from stored reports |
adaptive-thresholds.history-size / .min-history | 20 / 5 | Previous non-RED runs considered / needed |
adaptive-thresholds.green-sigma / .yellow-sigma | 2.0 / 3.0 | Band distance from the historical mean, in standard deviations |
adaptive-thresholds.min-spread | 0.005 | Lower limit for the standard deviation |
Report, persistence and resources
| Property | Default | Meaning |
|---|---|---|
persistence.collection / .database | comparison_reports / default | Where reports are stored |
persistence.retention | forever | How long reports are kept, e.g. 365d (TTL index) |
labels.<key> | – | Labels stored with every report, e.g. labels.environment=test |
max-examples | 20 | Example keys per category and per changed path |
max-value-examples | 3 | Before/after values per changed path; 0 disables them |
top-changed-paths | 50 | Changed paths listed in the report |
max-tracked-paths | 10000 | Distinct paths tracked per side; bounds memory |
batch-size | 1000 | Cursor batch size |
no-cursor-timeout | false | Keep idle server cursors alive on very large collections |
progress-interval | 10s | Progress logging and listener interval |
The connection itself uses Spring Boot's spring.mongodb.* properties.
Reading a report
Start at the top and work down: verdict, the rules that fired, then the metrics that explain them.
In the browser
--html=report.html writes the report as a single HTML page next to the JSON: verdict, hints with
ready-to-copy options, rules, changed paths with before/after values and the structure diff. The page makes no
external requests, so it works offline, as a CI artifact or as an e-mail attachment. For a stored report, use
java -jar ditto-cli.jar report --label=batchJobId=4711 --html=report.html; in your own code,
inject ReportHtml.
Already have a report.json? Open it in the report viewer: it renders the
file in your browser and uploads nothing. See an example report.
The JSON
{
"verdict": "YELLOW",
"rules": [
{ "rule": "keySimilarity", "level": "GREEN", "observed": 0.9950,
"reason": "0.9950 meets GREEN (>= 0.9900) (matched 19900, added 50, removed 50)" },
{ "rule": "maxPathChangeRate", "level": "YELLOW", "observed": 0.0812,
"reason": "Path 'price' changed in 0.0812 of matched documents, at or above the GREEN limit 0.0500 ...",
"details": [ "price: 0.0812 (1616 documents)" ] },
...
],
"keys": { "matched": 19900, "added": 50, "removed": 50, "keySimilarity": { "kind": "exact", "value": 0.995 } },
"content": { "unchanged": 18284, "changed": 1616, "unchangedRate": { "value": 0.9188 } },
"topChangedPaths": [ { "path": "price", "changedDocs": 1616, "expected": false,
"examples": [ { "type": "OBJECT_ID", "value": "{\"$oid\": \"6553f1212161972337cc2db4\"}" } ],
"valueExamples": [ { "baseline": [ "412.07" ], "candidate": [ "0" ], ... } ] } ],
"structure": { "newPaths": [], "missingPaths": [], "typeShifts": [], "presenceDeltas": [] },
"examples": { "changed": [...], "added": [...], "removed": [...] },
"run": { "durationMillis": 686, "baselineCount": 19950, "settings": { ... } },
"warnings": [],
"hints": [ { "kind": "EXPECTED_CHANGE", "path": "price",
"message": "price changed in 8% of matched documents. If this field is supposed to change …",
"property": "ditto.expected-change-paths=price", "cliOption": "--expected=price" } ]
}| If you see… | It usually means… |
|---|---|
Many removed, few added | A partial or aborted load |
Many removed and many added | Key values changed format ("123" vs 123) or a new ID scheme |
| One path changed in ~100% of documents | A mapping/format change in that field, or a technical field that should be ignored |
| Many paths at low rates | Normal data churn |
| Fields that change in most documents every run | Legitimate churn: declare them as expected-change-paths |
Anything in hints | A ready-to-paste property or CLI option for the next run |
| A type shift with unchanged content | Same values, different BSON type: consumers that read typed values may break |
warnings about the path cap | A map with dynamic keys: add it to wildcard-paths |
Example keys are relaxed Extended JSON, ready for mongosh:
db.products.find({_id: {"$oid": "6553f1212161972337cc2db4"}}). Look the same key up in both
collections to see the difference.
CLI reference
Short options map onto the request. Any --ditto.* or
--spring.mongodb.* property also works.
| Option | Meaning |
|---|---|
--uri, --db | Connection string (default mongodb://localhost:27017) and default database (default test) |
--collection | Compares <collection>_backup with <collection> |
--baseline, --candidate | Or name both collections explicitly |
--baseline-db, --candidate-db | Database per side, if not --db |
--key | Key field |
--ignore, --ordered, --wildcard, --expected, --redact | Comma-separated path lists |
--mode=auto|full|sample, --sample-size=n | Comparison mode (default auto) |
--matched-only | Compare only documents whose key exists on both sides |
--null-equals-missing, --mixed-key-types, --verdict-basis | As in the configuration table |
--label=key=value | Label stored with the report, e.g. a batch job id; repeatable |
--persist | Store the report in MongoDB |
--out=file | Also write the report to a file |
--html=file | Also write the report as a self-contained HTML page |
report --id=…, --label=…, --collection=… | Print a stored report (and with --html write it as a page); the newest one if several match. Exits with its verdict |
--help | Full usage |
Test-data generator
java -jar ditto-cli.jar generate writes a baseline and a modified copy of product-like
documents. They include nested objects, arrays of objects, a dynamic-key map, an ordered event list,
decimals, dates and nullable fields.
| Option | Default | Meaning |
|---|---|---|
--baseline / --candidate | demo_backup / demo | Collection names (dropped and recreated) |
--docs | 10000 | Baseline size |
--seed | 42 | Same seed, same data |
--touch-sync | true | Refresh meta.syncedAt in every candidate document, like a real sync |
--changes | – | Comma-separated kind[:path]:fraction, e.g. modify:price:0.05,drop-field:legacyCode:1,int-to-double:qty:0.5,shuffle-arrays:1,delete-docs:0.01,add-docs:0.01 |
Good to know
- Memory stays flat. FULL mode is a single streaming merge-join over both collections sorted by
key. There's no
$lookup, and nothing is loaded into memory as a whole. - Key order is verified. The Java key order reproduces MongoDB's sort order (simple collation), and a test checks it against a real server. Every key read must be strictly greater than the previous one, so duplicate keys abort the comparison instead of skewing the numbers.
- Index custom key fields.
_idis always indexed. Other key fields need an index, or the server sorts the whole collection; ditto warns about this. - Compare after the writes are done. Reads aren't a snapshot.
More detail in the README, including limitations and a sketch of a JobRunr integration with rollback.