// Engineering Log
Databases: Part 6 — MongoDB
Published on 2026-09-21
MongoDB — a document-oriented database. Instead of tables with rows it stores collections of documents in BSON format — a binary form of JSON with additional types (dates, binary data, decimal numbers). A document can contain nested objects and arrays, and documents in the same collection do not have to share the same set of fields.
License
MongoDB Community Server — versions released since October 16, 2018 — is distributed under the SSPL (Server Side Public License). It allows free use of MongoDB in your applications, including commercial ones, but requires disclosing the source code of the entire service wrapper if you provide MongoDB itself to third parties as a service. The Open Source Initiative does not recognize SSPL as an open license, so formally MongoDB is not open source. drivers for programming languages are distributed under Apache 2.0.
For a company that simply uses MongoDB as the database for its application, SSPL does not impose restrictions. If you specifically need an open license, there is FerretDB — described below.
Supported branches as of September 2026 are 7.0 and 8.0, the latest stable is 8.3. You should install the latest stable version: that is what the documentation recommends.
Document model
Example of an order document:
{
"_id": ObjectId("66f1c2a4e13b5a0012345678"),
"number": "2026-0915",
"customer": { "name": "ООО «Ромашка»", "inn": "7700000000" },
"items": [
{ "sku": "A-100", "qty": 2, "price": NumberDecimal("1500.00") },
{ "sku": "B-200", "qty": 1, "price": NumberDecimal("4200.00") }
],
"status": "paid",
"created_at": ISODate("2026-09-15T10:12:00Z")
}In a relational database the same data would occupy three tables — orders, customers, line items — and reading an order would require joins. In MongoDB an order is read in a single query. This is the main advantage of the document model: data that the application always reads and modifies together is stored together.
An operation on a single document in MongoDB is atomic. Therefore a well-designed schema where related data is embedded in one document often doesn’t require transactions at all.
When the document model is justified
- Data is naturally represented as nested objects: an order with line items, an article with content blocks, a profile with settings.
- Record structures differ: a product catalog where a phone and a sofa have different sets of attributes, logs of events from different sources.
- Schema changes frequently at an early product stage: a new field can be started without migrating the whole collection.
- Horizontal scalability is needed: sharding is built-in and distributes a collection across multiple servers.
When to choose a relational database instead
- Data is highly connected: accounting where documents, counterparties, warehouses, and entries reference each other. Joins in MongoDB are performed via $lookup in the aggregation pipeline, and with many relations this is more complex and slower than JOIN in PostgreSQL.
- Strict integrity is important: foreign keys, constraints, checks at the database level are strengths of SQL RDBMS.
- Many analytical queries over arbitrary fields: SQL is more convenient for this.
- You only need JSON fields in a regular database: PostgreSQL with jsonb type and GIN indexes covers many tasks for which MongoDB is chosen, while retaining all SQL capabilities.
A common mistake is choosing MongoDB because of “no schema.” The application still has a schema; it’s just not described in the database. When the code changes, documents of the old format remain in the collection, and the application must handle them.
Schema validation
To prevent the database from accepting documents with the wrong structure, MongoDB supports validation using JSON Schema:
db.createCollection("orders", {
validator: {
$jsonSchema: {
bsonType: "object",
required: ["number", "customer", "items", "status", "created_at"],
properties: {
number: { bsonType: "string" },
status: { enum: ["new", "paid", "shipped", "cancelled"] },
items: { bsonType: "array", minItems: 1 },
created_at: { bsonType: "date" }
}
}
},
validationLevel: "strict", // check all inserts and updates
validationAction: "error" // reject invalid documents; "warn" — only write to the log
})For an existing collection rules are added with the collMod command. Mode validationLevel: “moderate” checks only new documents and changes to already valid ones — this is convenient when gradually cleaning up old data.
Indexes
Without an index MongoDB scans the whole collection. Indexes can be single-field, compound, on array elements, text, geospatial, TTL (documents are removed automatically), or unique.
db.orders.createIndex({ status: 1, created_at: -1 })
db.sessions.createIndex({ created_at: 1 }, { expireAfterSeconds: 86400 }) // TTL: delete after a day
For compound indexes the MongoDB documentation has the ESR rule: first fields used for exact matches (Equality), then sort fields (Sort), then range fields (Range). Whether a query uses an index is shown by explain(“executionStats”).
Replica set
In production MongoDB is run not as a single server but as a replica set: one primary node accepts writes, secondaries receive copies of the data. If the primary fails, the remaining nodes automatically elect a new one. The recommended minimum is three nodes.
A replica set is needed not only for high availability. Multi-document transactions (since 4.0) and change streams work only in a replica set or a sharded cluster — they are not available on a standalone server. Therefore even for development it’s sometimes useful to run a single-node replica set.
Security
Since 3.6 MongoDB by default listens only on the local address, but password authentication is disabled by default. Databases exposed to the internet without authentication are routinely found by automated scanners: data is deleted and a ransom note is left.
Minimum for a production setup in mongod.conf:
net:
bindIp: 127.0.0.1,10.0.0.5 # local and internal address, not public
security:
authorization: enabledBefore enabling authorization you need to create an administrator, and for the application — a separate user with rights only to its database (role readWrite). Port 27017 should be closed by a firewall; as with Redis, when running in Docker it should not be published on all interfaces.
Backups
For small and medium databases the mongodump utility from the MongoDB Database Tools package (installed separately from the server) is suitable:
mongodump --uri="mongodb://backup:ПАРОЛЬ@127.0.0.1:27017/?authSource=admin" \
--oplog --gzip --archive=/backup/mongo-$(date +%F).archive.gz
mongorestore --gzip --archive=/backup/mongo-2026-09-21.archive.gz --oplogReplayThe –oplog parameter works only on a replica set node and includes changes made during the dump — resulting in a consistent copy as of the end time. For a sharded cluster you must stop the balancer (sh.stopBalancer()) and distributed transactions before dumping. For large databases mongodump is slow; they are copied via filesystem snapshots or specialized tools.
FerretDB — an open alternative
FerretDB is an Apache 2.0 licensed project that accepts MongoDB protocol requests and executes them in PostgreSQL with the DocumentDB extension. Applications connect to it with a regular MongoDB driver. The developers position it as a replacement for MongoDB 5.0 and newer in many cases, but compatibility is incomplete: before switching you need to check which commands and operators your application uses. The advantage is that data is stored in PostgreSQL, which is already well-known for administration and backups.
Advantages
- The document model matches object structures in code; related data is read with one query.
- Flexible schema can be augmented with JSON Schema validation when needed.
- Built-in replication with automatic failover.
- Sharding to distribute data across servers.
- Mature query and aggregation language: filtering, projections, aggregation pipelines, full-text and geosearch.
Disadvantages
- SSPL license is not recognized as open.
- Data joins via $lookup are more complex and slower than relational JOINs.
- Transactions are available only in a replica set and put noticeable load on the database; they should be used sparingly.
- Data size is usually larger than in a relational database: field names are stored in every document. The size of a single document is limited to 16 MB.
- Cluster administration: a replica set and especially a sharded cluster are harder to operate than a single PostgreSQL server.
MongoDB is well suited for data that naturally fits into documents and changes structure, and for systems that need horizontal scalability. For accounting systems with many relations, strict integrity, and SQL analytics, it is safer to choose PostgreSQL — possibly with jsonb fields for the parts of data that truly need flexibility.
// Contact
Need help?
Get in touch with me and I'll help solve the problem
I reply within one business day (03:00-13:00 GMT)
Или оставьте заявку здесь:
// Related