Change-aware upserts for MongoDB & Spring Data
Write every document. Emit only real changes.
Diffsert writes your documents so that MongoDB records only the fields that changed.
Change streams and Debezium see a small update event instead of a full
replace, and documents that didn't change produce no event at all.
- Maven Centralv0.1.0
- MongoDB 5.0 – 9.0
- Spring Boot 4 starter
- Java 25
- MIT
{ operationType: "replace",
fullDocument: {
_id: "c1", name: "Acme", city: "Hamburg",
email: "billing@acme.example", tier: "gold",
jobRunId: "run-2",
_createdAt: ISODate("2026-10-08T12:00Z"),
_updatedAt: ISODate("2026-10-08T12:00Z") } }
{ operationType: "update",
updateDescription: {
updatedFields: {
city: "Hamburg",
jobRunId: "run-2",
_updatedAt: ISODate("2026-10-08T12:00Z") },
removedFields: [] } }
jobRunId and _updatedAt? No write, no event.One batch run, 1,000 documents
Your consumers should hear about 15 documents, not 1,000.
A sync job maps rows from another system to DTOs and writes them in batches. With save(),
every write replaces the whole document, so every document looks changed on every run. Run metadata such as
jobRunId makes it worse: it differs every time. Diffsert compares before it writes.
save() / replaceOne
- 1,000 replace
Diffsert upsertAll
- 3 insert
- 12 update
- 985 unchanged, no event
Example run. The result object reports the same split:
requested=1000, inserted=3, updated=12, unchanged=985, notFound=0
| Situation | save() / replaceOne | Diffsert |
|---|---|---|
| Document is new | insert event | insert event |
| Nothing changed | replace event, full document | No write, no event |
| Only run metadata changed | replace event, full document | No write, no event (ignored fields) |
city changed | replace event, full document | update event with city and the ignored fields |
| A field was removed | replace event, full document | update event, removedFields: ["name"] |
_createdAt in the new document | Overwritten | Stored value kept (preserved field) |
How it works
A pipeline update instead of a replace.
Every write is an updateOne by _id with an aggregation pipeline. The stored result
is the same as with a replace, but MongoDB 5.0+ diffs it against the stored document and writes only the
difference to the oplog.
-
Wrap the new document in
$literalValues starting with
$, such as"$5 off", are written as text and never evaluated as expressions. -
Copy back preserved fields
Fields like
_createdAtkeep their stored value once set. On the first insert, the new value is used. -
Compare without the ignored fields
If only
jobRunIdor_updatedAtdiffer, the pipeline returns the stored document unchanged. MongoDB sees a no-op: matched, not modified, no event. -
MongoDB logs the delta
For a real change, the server writes a
$v: 2delta entry to the oplog. Change streams turn it into anupdateevent with just the changed fields.
[{ $replaceWith: { $let: {
vars: { incoming: { $literal: <new document> } },
in: { $let: {
vars: { result: <incoming + stored _createdAt> },
in: { $cond: {
if: { $and: [ <document existed>,
{ $eq: [ <$$ROOT minus ignored fields>,
<$$result minus ignored fields> ] } ] },
then: "$$ROOT", // nothing relevant changed: no-op
else: "$$result"
} } } } } } }]
{ op: "u", ns: "app.customers", o2: { _id: "c1" },
o: { $v: 2, diff: { u: {
city: "Hamburg", jobRunId: "run-2",
_updatedAt: ISODate("2026-10-08T12:00Z") } } } }
Quick start
Pick your level of Spring.
Three modules, one mechanism, all on Maven Central under dev.jbaby.
<dependency>
<groupId>dev.jbaby</groupId>
<artifactId>diffsert-spring-boot-starter</artifactId>
<version>0.1.0</version>
</dependency>
diffsert:
ignore-fields: jobRunId, _updatedAt # alone, these are not a change
preserve-fields: _createdAt # stored value wins once set
@Autowired Diffsert diffsert;
BatchWriteResult result = diffsert.upsertAll(dtos, CustomerDto.class);
// requested=1000, inserted=3, updated=12, unchanged=985, notFound=0
WriteOutcome outcome = diffsert.upsert(dto);
// INSERTED / UPDATED / UNCHANGED / NOT_FOUND
The starter creates a Diffsert bean from the application's MongoTemplate and the diffsert.* properties.
- With Micrometer, written documents are counted as
diffsert.documents, tagged withcollectionandoutcome. - Every
DiffsertListenerbean is called after each write. A failing listener is logged and doesn't affect the write. - To configure options in code, define a
DiffsertOptionsbean; it wins over the properties. - Your own
Diffsertbean replaces the auto-configured one, including the listener wiring. Attach metrics withdiffsert.withListener(diffsertMetrics).
<dependency>
<groupId>dev.jbaby</groupId>
<artifactId>diffsert-spring-data</artifactId>
<version>0.1.0</version>
</dependency>
Diffsert diffsert = new Diffsert(mongoTemplate, DiffsertOptions.builder()
.ignoreForChangeDetection("jobRunId", "_updatedAt")
.preserveExistingValue("_createdAt")
.build());
diffsert.upsertAll(dtos, CustomerDto.class); // collection from @Document
diffsert.upsertAll(dtos, "customers"); // or explicit
Entities are converted the way save() converts them, with Spring Data's MongoConverter: @Id, @Field and custom converters apply. Every entity needs an id.
- The type key
_classis removed by default;withTypeKeyRemoved(false)keeps it. updateandupdateAllnever insert.toWriteModels(...)returns the write models without executing them, for sessions, transactions or your ownbulkWrite.
<dependency>
<groupId>dev.jbaby</groupId>
<artifactId>diffsert-core</artifactId>
<version>0.1.0</version>
</dependency>
DiffsertWriter writer = new DiffsertWriter(DiffsertOptions.builder()
.ignoreForChangeDetection("jobRunId", "_updatedAt")
.build());
MongoCollection<Document> customers = database.getCollection("customers");
BatchWriteResult result = writer.upsertAll(customers, documents);
diffsert-core depends only on the MongoDB Java driver. It takes plain Documents with an _id.
- The collection is passed per call, so its codecs and write concern apply as you configured them.
buildPipeline(document)returns the update pipeline itself, if you want to run it your own way.withListener(...)reports every write, including the succeeded part of a failed bulk write.
Options
Five settings, sensible defaults.
| Builder | Property | Default | Meaning |
|---|---|---|---|
ignoreForChangeDetection(…) | diffsert.ignore-fields | none | Top-level fields that alone don't count as a change. They are written along with real changes. |
preserveExistingValue(…) | diffsert.preserve-fields | none | Top-level fields whose stored value is kept once set, such as _createdAt. |
replace() / merge() | diffsert.mode | replace | replace removes fields missing from the new document. merge keeps fields that only the stored document has. |
ordered(…) | diffsert.ordered | false | Ordered bulk writes stop at the first error. Unordered ones continue and are faster. |
withTypeKeyRemoved(…) | diffsert.remove-type-key | true | Remove Spring Data's type key _class from converted entities. |
Modules & versions
Tested against every MongoDB line it supports.
The delta oplog behavior Diffsert relies on is a server implementation detail, not a documented guarantee. So CI runs every test against each supported MongoDB version, on every change and once a week.
- 5.0
- 6.0
- 7.0
- 8.0
- 8.2
- 8.3
- 9.0
Auto-configured Diffsert bean from diffsert.* properties, listener beans and Micrometer counters.
Entities in, converted with MongoConverter. For Spring applications without Boot auto-configuration.
Documents in, MongoDB Java driver only. The pipeline, options and results live here.
MongoDB 5.0 or later
Change streams and Debezium need a replica set. The writes themselves also work on a standalone server.
Java 25
Built and tested with Java 25.
Spring Boot 4
The Spring modules are built against Spring Boot 4.0, Spring Data MongoDB 2025.1 and driver 5.6.
Caveats
What to know before you rely on it.
Each of these is checked by an integration test on every supported MongoDB version.
-
Small documents can still produce
replaceeventsMongoDB logs a delta only when it is smaller than the new document. For very small documents, or when almost every field changes, change streams emit a
replaceevent. In a test with ten short fields, changing nine still gave anupdate. Consumers should handle both. -
Unchanged documents are not touched
They keep their old run metadata, such as
jobRunId. "Not written in this run" therefore can't detect records deleted at the source. Track the ids you saw separately. -
Keep field order and types stable
A different field order counts as a change, for example after reordering DTO fields. A type change alone, such as
1to1L, is written without ignored or preserved fields; with them, equal numbers count as unchanged and the stored type stays. -
Arrays and nested documents are reported by path
A changed element appears as
tags.2, a nested field asaddress.street. Shortened arrays appear intruncatedArrays. -
Top-level fields, matched by
_idIgnored and preserved fields must be top-level. Documents are matched by
_idonly, and every document needs one. -
Driver exceptions are not translated
A
MongoBulkWriteExceptionreaches you as-is, withgetWriteErrors(). Documents written before or besides the failed ones are still reported to listeners and metrics.