How to Model Relationships in NoSQL Databases
DEV Community

How to Model Relationships in NoSQL Databases

After you move from a SQL database to NoSQL, you usually forget about: - Joins - Foreign keys - Normalization and Denormalization - Database migrations Yet, most developers have no idea how to model relationships in NoSQL correctly. Most developers come to MongoDB (for example) after years of SQL. So they model the data the way SQL taught them. Customers, Orders and Order Items - all go in multiple collections by habit. Every link becomes an ID, just like a foreign key. Then they write the joins by hand in C#, because MongoDB has no JOIN keyword. The result is a document database that behaves like a slow relational one. A single read needs four database round trips. I've run MongoDB in production for years. Almost every performance problem I've seen there came from the wrong data model. Today you will learn how to model relationships in a document database. In this post, we will explore: - Why Relationships Work Differently in NoSQL - Embed or Reference: The Core Decision - One-to-One: Embed by Default - One-to-Many: The Three Sizes - One-to-Few: Embed the Array - One-to-Many: Reference from the Child - One-to-Zillions: Never Grow an Array Forever - Many-to-Many: Orders and Products - Reading Related Data Without Joins - Keeping Duplicated Data Correct - 4 Mistakes That Break NoSQL Data Models - The Same Rules in DynamoDB, Cassandra and Cosmos DB All the code in this article I show using MongoDB and C#, but the rules apply to any NoSQL database. Let's dive in. ๐Ÿ‘‰ Read original article on my newsletter: https://antondevtips.com/blog/how-to-model-relationships-in-nosql-databases Why Relationships Work Differently in NoSQL In SQL databases, you have tables and rows. In NoSQL databases, you have collections and documents. A document is one record. It's a JSON-like object that can contain other objects and arrays. A collection is a group of documents. It's the closest thing MongoDB has to a table. MongoDB has no foreign key constraints. Nothing stops you from deleting a customer who still has orders. There are no cascade deletes either. Delete a parent, and the children stay behind as orphans. And a normal query can't join. MongoDB can join inside the aggregation pipeline, but that isn't how you read data most of the time. In SQL you design the tables first and write the queries later. You normalize the data, which means you store each fact once and link to it. The query planner works out the rest. In a document database, you do the opposite. You shape the documents around the queries you already know you need. Here's the same shipment in both worlds. In SQL it's three tables: shipments (id, number, order_id, status) shipment_address (shipment_id, street, city, zip) shipment_items (id, shipment_id, product, quantity) Reading one shipment means joining all three. In MongoDB, it's one document: { "_id": "8e3b1a6c-2f45-4d17-9b0a-51c9e2f7a410", "number": "12345678", "orderId": "ORD-9931", "status": "Dispatched", "address": { "street": "Main St 15", "city": "Warsaw", "zip": "00-001" }, "items": [ { "product": "Laptop", "quantity": 1 }, { "product": "Mouse", "quantity": 2 } ] } One read returns the whole shipment, without joins. Side by side, the two shapes look like this: SQL: 3 tables MongoDB: 1 document ------------------------------------------- --------------------------------- shipments (id, number, status) shipment { number, status, shipment_address (shipment_id, street, address: { street, city, zip }, city, zip) items: [ { product, quantity } ] shipment_items (id, shipment_id, } product, quantity) Embed or Reference: The Core Decision In a NoSQL database, you have just two tools. Embed means you put the related data inside the parent document. The address above is embedded in the shipment. Reference means you store an id and keep the data in its own collection. That's the same idea as a foreign key, except the database doesn't enforce it. Every relationship in your model is one of those two. Picking the right one is most of the work. Here's the rule I use: | Embed when | Reference when | |---|---| | You read both together almost every time | The child is often read on its own | | The list has a known upper limit | The list can grow infinitely | | The child changes when the parent changes | The child is updated on its own schedule | | The child belongs to one parent | Many parents share the same child | | The data is small | The data is large | One hard limit settles a lot of arguments. A MongoDB document can't be bigger than 16 MB. If a list can outgrow that, it can't be embedded, and there's nothing left to discuss. Speed pushes the other way. Embedded data comes back in the same read, so it costs nothing extra to fetch. Note: A good model embeds some things and references others. Don't pick one style and apply it everywhere. Now I will show you the three shapes a relationship can take. ๐Ÿ‘‰ Read original article on my newsletter: https://antondevtips.com/blog/how-to-model-relationships-in-nosql-databases Top comments (0)

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.