Ensiklopedia VibeKoding: Data Models: A Complete Overview โ Document, Graph, Time-Series, and Vector.Ensiklopedia VibeKoding: Data Models: A Complete Overview โ Document, Graph, Time-Series, and Vector.
Why can't you just stuff all your data into MySQL tables? When your data is a social network graph, millions of sensor readings per second, or semantic vectors for AI to understand, relational tables fall short. Different data shapes require different modeling approaches.Why can't you just stuff all your data into MySQL tables? When your data is a social network graph, millions of sensor readings per second, or semantic vectors for AI to understand, relational tables fall short. Different data shapes require different modeling approaches.
------
Relational databases (MySQL, PostgreSQL) organize data with "tables + rows + columns," suitable for structured, well-defined business data. But real-world data comes in far more forms than just this:Relational databases (MySQL, PostgreSQL) organize data with "tables + rows + columns," suitable for structured, well-defined business data. But real-world data comes in far more forms than just this:
| Data Shape | Relational Pain Point | Better Model |
|---|---|---|
| User profiles (flexible fields, nested structures) | Frequent ALTER TABLE, many NULL columns | Document Model |
| Social networks (friends of friends of friends) | Multi-level JOIN performance degrades exponentially | Graph Model |
| Monitoring metrics (millions of writes per second) | Write bottlenecks, historical data bloat | Time-Series Model |
| AI semantic search ("similar meaning" content) | Cannot express semantic similarity | Vector Model |
It's not about "replacing" relational databases, but "supplementing" them. Most systems still run their core business on MySQL/PostgreSQL, but introducing specialized data models for specific scenarios can yield orders-of-magnitude performance improvements.It's not about "replacing" relational databases, but "supplementing" them. Most systems still run their core business on MySQL/PostgreSQL, but introducing specialized data models for specific scenarios can yield orders-of-magnitude performance improvements.
------
The document model stores data as JSON/BSON documents, where each record is a self-contained document that can have different field structures.The document model stores data as JSON/BSON documents, where each record is a self-contained document that can have different field structures.
json { "_id": "user_1001", "name": "Zhang San", "tags": ["VIP", "Active"], "address": { "city": "Beijing", "district": "Chaoyang" }, "orders": [ { "id": "o1", "amount": 299 }, { "id": "o2", "amount": 599 } ] }
Key Features:Key Features:
| Comparison | Relational (MySQL) | Document (MongoDB) |
|---|---|---|
| Data Structure | Fixed Schema, ALTER TABLE to modify | Flexible Schema, add fields anytime |
| Nested Data | Requires multi-table JOINs | Embedded directly in the document |
| Cross-record Relationships | JOINs are powerful | Relationship queries are weaker |
| Best For | Structurally stable business data | Structurally variable content data |
"MongoDB doesn't need data structure design" โ Wrong! The document model also requires careful design: nesting levels shouldn't be too deep, and frequently updated sub-documents should be split into separate collections."MongoDB doesn't need data structure design" โ Wrong! The document model also requires careful design: nesting levels shouldn't be too deep, and frequently updated sub-documents should be split into separate collections.
------
The graph model uses Nodes and Edges to represent entities and their relationships. Each node is an entity, each edge is a relationship, and both nodes and edges can carry properties.The graph model uses Nodes and Edges to represent entities and their relationships. Each node is an entity, each edge is a relationship, and both nodes and edges can carry properties.
CODE (Zhang San) --[follows]--> (Li Si) --[follows]--> (Wang Wu) | | +--------[purchased]----> (iPhone) <--[purchased]--+
Scenario: Finding "friends of friends of friends" in a social networkScenario: Finding "friends of friends of friends" in a social network
Relational approach (3-level JOIN):Relational approach (3-level JOIN):
sql SELECT DISTINCT f3.name FROM friends f1 JOIN friends f2 ON f1.friend_id = f2.user_id JOIN friends f3 ON f2.friend_id = f3.user_id WHERE f1.user_id = 1001;
Graph database approach (Cypher query language):Graph database approach (Cypher query language):
cypher MATCH (me)-[:FOLLOWS*1..3]->(target) WHERE me.name = 'Zhang San' RETURN DISTINCT target.name
Each additional hop in the relational approach adds another JOIN, causing exponential performance degradation. Graph databases traverse relationships via pointers directly, so multi-hop query performance remains nearly unchanged.Each additional hop in the relational approach adds another JOIN, causing exponential performance degradation. Graph databases traverse relationships via pointers directly, so multi-hop query performance remains nearly unchanged.
------
The time-series model uses timestamps as the primary axis, specifically optimized for "write in chronological order, query by time range" scenarios.The time-series model uses timestamps as the primary axis, specifically optimized for "write in chronological order, query by time range" scenarios.
CODE timestamp device cpu_usage memory 2024-01-15 10:00:01 server-01 45% 12.3GB 2024-01-15 10:00:02 server-01 67% 12.5GB 2024-01-15 10:00:03 server-01 92% 14.1GB
| Issue | MySQL | Time-Series Database (InfluxDB) |
|---|---|---|
| Write Speed | Tens of thousands/sec | Millions/sec |
| Historical Data | Manual cleanup, tables keep growing | Automatic expiration policy (TTL) |
| Aggregation Queries | Slow GROUP BY | Built-in downsampling (5 sec โ 1 min average) |
| Storage Efficiency | General-purpose storage, wasted space | Columnar compression, saving 90% space |
------
The vector model converts unstructured data like text, images, and audio into high-dimensional numerical vectors through an Embedding model, then measures semantic similarity by calculating the distance between vectors.The vector model converts unstructured data like text, images, and audio into high-dimensional numerical vectors through an Embedding model, then measures semantic similarity by calculating the distance between vectors.
CODE "delicious Japanese food" โ Embedding โ [0.82, 0.15, 0.91, 0.33, ...] โ Cosine similarity "Ginza sushi master" โ [0.80, 0.18, 0.89, ...] โ 96% similar "Italian pizza" โ [0.12, 0.85, 0.20, ...] โ 31% similar
| Comparison | Keyword Search (LIKE / Full-text Index) | Vector Search |
|---|---|---|
| Search Method | Exact string matching | Semantic similarity matching |
| "delicious Japanese food" | Can only match text containing "Japanese food" | Can find "sushi," "sashimi," "izakaya" |
| Multilingual | Needs separate handling | Cross-language semantic understanding |
| Multimodal | Text only | Unified retrieval across text, images, and audio |
- Standalone Vector Databases: Pinecone, Milvus, Weaviate โ focused on vector retrieval, best performance - Traditional Database Extensions: pgvector (PostgreSQL), Atlas Vector Search (MongoDB) โ reduce architectural complexity - In-Memory Vector Libraries: FAISS, Annoy โ suitable for small-scale, low-latency scenarios- Standalone Vector Databases: Pinecone, Milvus, Weaviate โ focused on vector retrieval, best performance - Traditional Database Extensions: pgvector (PostgreSQL), Atlas Vector Search (MongoDB) โ reduce architectural complexity - In-Memory Vector Libraries: FAISS, Annoy โ suitable for small-scale, low-latency scenarios
------
| What Does Your Data Look Like? | Recommended Model | Representative Products |
|---|---|---|
| Fixed structure, clear relationships (orders, users) | Relational | MySQL, PostgreSQL |
| Flexible structure, deep nesting (content, configs) | Document | MongoDB, DynamoDB |
| Complex relationships between entities, need multi-hop traversal | Graph | Neo4j, Amazon Neptune |
| Write in chronological order, query by time range | Time-Series | InfluxDB, TimescaleDB |
| Unstructured data, need semantic similarity search | Vector | Pinecone, Milvus, pgvector |
Modern systems typically use multiple models together: - Core business on PostgreSQL (relational) - User behavior logs on InfluxDB (time-series) - AI knowledge base on Milvus + pgvector (vector) - Recommendation engine on Neo4j (graph) Don't try to find "one database to solve all problems" โ instead, let each type of data find its most suitable home.Modern systems typically use multiple models together: - Core business on PostgreSQL (relational) - User behavior logs on InfluxDB (time-series) - AI knowledge base on Milvus + pgvector (vector) - Recommendation engine on Neo4j (graph) Don't try to find "one database to solve all problems" โ instead, let each type of data find its most suitable home.