How to Build an Open Lakehouse on Your Laptop
DEV Community

How to Build an Open Lakehouse on Your Laptop

open-lakehouse is a small, complete lakehouse built only from open-source parts: object storage, an open table format, a catalog, Spark as the engine, a stream, a pipeline framework, an orchestrator and experiment tracking. It runs on Docker Compose and is driven by one CLI, ./lakehouse . This post builds it on your machine one layer at a time, in the order the layers depend on each other, and checks each layer before the next one goes on top of it, which is the whole trick to a stack like this: when something breaks later, you already know everything underneath it works. Each section covers what the layer is and why a lakehouse needs it, the exact commands, and a check with the real output from my run (trimmed where it was long). It ends with a small diagram of the stack so far, with the new piece in orange and what you already built in grey. What you're building | Layer | Component | Version | Ports | |---|---|---|---| | Storage | SeaweedFS (S3 API); PostgreSQL | 3.80; 16.15 | 8333; 5432 | | Compute | Apache Spark standalone cluster + Spark Connect server | 4.2.0 | 7078 (UI 8082, worker UI 8083); 15002 | | Table format | Delta Lake | 4.4.0 | | | Catalog | Unity Catalog OSS, plus its web UI | 0.6.0 | 8081; 3001 | | Pipelines | Spark Declarative Pipelines | ships with Spark | | | Streaming | Structured Streaming, including Real-Time Mode | ships with Spark | | | Event log | Kafka (Confluent Platform 7.9) + ZooKeeper | 3.9 | 9092; 2181 | | Orchestration | Apache Airflow | 3.3.2 | 8085 | | ML | MLflow tracking (and AI Gateway) | 3.16.1 | 5000 (5001) | | Optional | Delta Sharing server; read-only dashboard | 1.3.10 | 8443; 3000 | Everything runs on one Spark version, 4.2.0. The Spark images, the Python clients (your terminal's, Airflow's and Jupyter's) and the JARs that plug Delta, Unity Catalog, S3 and Kafka into Spark are all pinned to match it, and the repo's README has the full versions table with the reasoning for each pin. Every service runs in Docker on one Compose bridge network, including storage, so the host needs only the tools that drive it: - Docker with the Compose plugin, version 2.30 or newer (the Spark compose file uses post_start hooks) - Python 3.10+ and Poetry 2.x - curl ,git , and ideallync Plan for about 16 GB of RAM and 20 GB of disk. Spark asks for a 4 GB driver and 8 GB executors by default; on a smaller machine, lower spark.executor.memory in config/spark/spark-defaults.conf once setup has created it. Linux first, with Mac and Windows notes I ran every command below on Ubuntu 24.04 with Docker Engine and the Compose plugin, and the steps assume Linux. - macOS. Use Docker Desktop, plus brew install python@3.12 poetry coreutils netcat git (setup checks forgdf , the GNUdf ). Under Settings > Resources, give Docker Desktop enough memory for Spark (16 GB is comfortable). You don't need Docker Desktop's host-networking setting: every service sits on the Compose bridge network and publishes its ports. Keep the repo under your home directory, which Docker Desktop shares by default, because the Spark containers bind-mount./data ,./jars and./config . The repo's installation guide has the details; I haven't run this walkthrough on a Mac myself. - Windows. Use WSL2 with Ubuntu and Docker Desktop's WSL integration, then follow the Linux steps inside WSL. Clone the repo into your WSL home directory, not under /mnt/c , for the same bind mounts. Get the code: git clone https://github.com/open-lakehouse/open-lakehouse.git cd open-lakehouse If you already have a clone from an earlier version of this post, run git pull in it instead. The repo pins every image and JAR to a combination that was tested together, so the newest commit is the one to follow along with, and the next section brings your images and JARs up to it. Setup: config files, JARs, Python deps and images Everything goes through ./lakehouse , a bash script at the root of the repo that wraps Docker Compose, checks ports before it starts anything, and knows which compose file and containers belong to each layer. ./lakehouse help lists its commands. The first setup creates the two config files you edit, .env and config/spark/spark-defaults.conf , from their checked-in examples, and stops: ./lakehouse setup 2. Checking configuration files... ! Created .env from .env.example → Please edit .env with your credentials ! Created spark-defaults.conf from example → It uses the demo S3 pair from .env.example; keep the two files in step ... Setup incomplete - please address the issues above Both files are gitignored. The S3 keys in them are already a matching local demo pair (lakehouse_s3 / lakehouse_s3_secret ), the same pair the SeaweedFS container is created with, so the only values you have to set are the PostgreSQL login and the Airflow admin login: sed -i 's/^POSTGRES_USER=./POSTGRES_USER=lakehouse/; s/^POSTGRES_PASSWORD=./POSTGRES_PASSWORD=lakehouse-local/; s/^AIRFLOW_ADMIN_USER=./AIRFLOW_ADMIN_USER=admin/; s/^AIRFLOW_ADMIN_PASSWORD=./AIRFLOW_ADMIN_PASSWORD=admin-local/' .env (On macOS, BSD sed wants sed -i '' .) Pick your own passwords; these are only local. If you ever change the S3 keys, change them in .env , spark-defaults.conf and config/unity-catalog/server.properties together, and ./lakehouse check-config compares them for you. Run setup again. This time it checks your tools, runs poetry install (which brings PySpark 4.2.0), downloads any JAR the Spark containers load that's missing from ./jars (about 700 MB the first time, mounted at /opt/spark/jars-extra ), and creates ./data writable by both you and the Spark containers: ./lakehouse setup 2. Checking configuration files... ✓ .env exists ✓ Required variables set ✓ spark-defaults.conf exists ✓ Credentials consistent 3. Checking JAR dependencies... ✓ All required JARs present 4. Installing Python dependencies... ✓ Python dependencies installed 5. Preparing data directory and database bootstrap... ✓ ./data ready (shared by host test-data tools and the Spark containers) ✓ PostgreSQL runs in Docker; its databases are created by './lakehouse start storage' ... Setup complete! Last, pull the images. ./lakehouse pull pulls every image the compose files pin; the few images the repo builds itself (Airflow, MLflow, the dashboard, Delta Sharing) are rebuilt on their first start, and again whenever a git pull changes them: ./lakehouse pull storage (docker-compose-storage.yml) spark (docker-compose-spark.yml) kafka (docker-compose-kafka.yml) ... ✓ Registry images match the docker-compose-*.yml pins Locally built images (airflow, mlflow, notebooks, dashboard, delta-sharing) rebuild on their next start. The JARs are the lakehouse's plumbing, and their versions have to agree with each other and with Spark: | JAR | Version | What it's for | |---|---|---| delta-spark_4.2_2.13 , delta-storage | 4.4.0 | Delta Lake, built for Spark 4.2 | unitycatalog-spark_4.2_2.13 , -client , -hadoop | 0.6.0 | Spark's connector to Unity Catalog, matching the server | hadoop-aws , AWS SDK v2 bundle , analyticsaccelerator-s3 | 3.5.0, 2.35.4, 1.3.1 | the s3a filesystem; must match the Hadoop inside the Spark image | spark-sql-kafka-0-10_2.13 and friends | 4.2.0 | the Kafka source and sink for Structured Streaming | Storage: SeaweedFS and PostgreSQL A lakehouse keeps its tables as ordinary files in object storage: Parquet data files plus the table format's own log. SeaweedFS is a small open-source object store that speaks the S3 API, so Spark writes to it exactly the way it would write to Amazon S3, through Hadoop's s3a connector. PostgreSQL sits beside it as the metadata database for the services that need one (Airflow and MLflow). ./lakehouse start storage start storage brings up both containers, waits until Postgres actually accepts connections, then creates the lakehouse bucket and a few prefixes: Starting storage (SeaweedFS + PostgreSQL)... Container postgres Started Container seaweedfs Started ✓ PostgreSQL ready ✓ SeaweedFS S3 ready Bootstrapping the bucket and warehouse prefixes... SeaweedFS S3 (http://localhost:8333, bucket lakehouse): ✓ created bucket lakehouse ✓ prefix s3://lakehouse/warehouse/bronze/ ... PostgreSQL (container postgres): ✓ accepts connections as lakehouse (airflow and mlflow create their databases on first start) storage bootstrap complete. Storage ready. To check it, list the bucket. SeaweedFS enforces the S3 credentials, so sign the request; curl can do that by itself: curl -s --aws-sigv4 "aws:amz:us-east-1:s3" --user lakehouse_s3:lakehouse_s3_secret \ "http://localhost:8333/lakehouse?list-type=2" | grep -o ' [^ ' warehouse/_checkpoints/ warehouse/bronze/ warehouse/gold/ warehouse/pipeline-history/ warehouse/silver/ flowchart LR classDef new fill:#d97706,fill-opacity:0.72,stroke:#d97706,color:#ffffff classDef old fill:#4b5563,fill-opacity:0.72,stroke:#4b5563,color:#ffffff subgraph d1_ST["storage"] d1_SW[("SeaweedFS :8333 bucket lakehouse")]:::new d1_PG[("PostgreSQL :5432")]:::new end Compute: Spark 4.2 and Spark Connect Spark is the engine: it reads and writes every table, and everything above this layer either feeds it or tells it what to run. The stack runs a small standalone cluster (one master, one worker) plus a Spark Connect server. Connect splits the classic Spark application in two: the server owns the driver, the catalog settings and the storage credentials, and clients send it query plans over gRPC from anywhere, as thin Python processes with no JVM. Every client in this post, from your terminal to Airflow, talks to sc://localhost:15002 . ./lakehouse start spark Starting Spark 4.2.0 (port 7078, Connect on 15002)... ✓ Spark Connect ready Spark Connect endpoint: sc://localhost:15002 ✓ Spark Master ready ✓ Spark UI ready "Spark Connect ready" means the server is listening inside its container, not merely that Docker has published the port. The Connect server is one long-running Spark application sharing the worker with any job you submit straight to the cluster, so it's set up not to crowd

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.