Optimizing Boot Time: Techniques to Reduce Time-to-Shell
DEV Community

Optimizing Boot Time: Techniques to Reduce Time-to-Shell

  • Measuring the boot path and exposing the real hotspots - Squeeze the earliest seconds: practical SPL, DTB and U‑Boot tuning - Make the kernel and initramfs faster: compression, initcalls and modules - Service ordering and filesystem tricks that shave seconds - Practical application: checklists and recipes to cut seconds from boot Boot time is an engineering problem you solve with measurement, not magic. In my board bring‑up work a single mis‑configured SPL or an over‑verbose bootloader has routinely eaten multiple seconds between power and a usable shell - and those seconds add up across thousands of devices and test cycles. The symptom is always the same: board teams report “slow boot” and we see a scatter of effects - long SPL/DRAM init, U‑Boot autoscans, big kernel decompress, or a userspace service blocking for network. Those hold-ups translate to longer R&D iterations, slower factory test throughput and lower perceived quality in the field. The first rule: you must measure the entire chain (hardware toggles through kernel traces and userspace timelines) and isolate the single longest path before changing knobs. Measuring the boot path and exposing the real hotspots Accurate measurement wins the argument and prevents wasted optimization work. Use a mix of hardware and software telemetry so you can attribute every millisecond. - Hardware boundary markers - Toggle a dedicated GPIO in SPL, in U‑Boot right before handover, and in the kernel early init to get wall‑clock boundaries with an oscilloscope or logic analyzer. This gives an unambiguous timeline from reset to kernel handoff and to init. Hardware toggles avoid any logging‑related distortion. - Bootloader and kernel prints - Enable earlyprintk and kernel timestamping withprintk.time=1 to get kernel-side timestamps in the logs. These parameters are documented in the kernel command‑line reference. - Use initcall_debug on the kernel command line to print per‑initcall durations; that exposes slow static driver init work. - Enable - Kernel tracing for deep dives - Use ftrace viatrace-cmd / KernelShark to capture fine‑grained boot events and visualize CPU‑side hotspots. This uncovers driver probe stalls and IRQ/lock contention during early init. - Use - Userspace timelines - With systemd usesystemd-analyze time ,systemd-analyze blame andsystemd-analyze critical-chain to split the boot into kernel / initramfs / userspace and identify long services.systemd-analyze plot generates an SVG flame-chart of service startup order. - With - Persistent, cross‑reboot logs - Configure pstore /ramoops to persist early kernel logs or ftrace across reboots so you don’t lose the data on a crash during experiments. - Configure Example quick checklist to gather data: # 1) U-Boot: reduce autoboot while you instrument: setenv bootdelay 3 # 2) Kernel command line (temporary testing): console=ttyS0,115200 earlyprintk=serial,ttyS0,115200 printk.time=1 initcall_debug # 3) Capture userspace timing after boot: systemd-analyze time systemd-analyze blame > /tmp/boot-blame.txt systemd-analyze critical-chain > /tmp/critical-chain.txt # 4) For function-level traces: trace-cmd record -e boot -o /tmp/boot.dat -- Cite the standard tooling and parameters when you automate this measurement. Important: measurement must be repeatable. Automate a harness (power cycling with a relay) and collect many samples; statistical outliers often point to hardware readiness race conditions. Squeeze the earliest seconds: practical SPL, DTB and U‑Boot tuning The earliest few seconds are won in the SPL/U‑Boot space. SPL exists to do as little as possible and hand off to U‑Boot (or directly to firmware). Make it minimal and deterministic. The U‑Boot project documents the SPL build model and the knobs you should trim. What to do in SPL - Build only what SPL absolutely needs: DRAM init, minimal console (optionally disabled in production), power rails, and the loader for your payload. Remove filesystem drivers, splash logic and non‑essential hardware services from SPL. The SPL build supports explicit CONFIG_SPL_* toggles to reduce the object set. - Use a smaller, filtered DTB in SPL. U‑Boot’s SPL build uses fdtgrep to produce a much smaller SPL DTB - strip nodes not required before RAM relocation. - Avoid dynamic hardware enumeration during SPL. Hardcode timings and DDR settings for production‑grade boards once DDR training is validated; dynamic training is useful during bring‑up but costs time. U‑Boot configuration and environment - Set the environment to production defaults: bootdelay=0 ,autoload=no , and a deterministicbootcmd . Avoid menus and interactive timeouts in production. - Keep console output minimal during production boots: use silent_linux or setbootargs so kernel prints are reduced to the minimalloglevel . Excessive console prints (serial/console I/O) can cost hundreds of milliseconds to seconds on slow UARTs. - Bundle kernel, DTB and optional initramfs as a FIT image and boot a single image blob rather than doing multiple loads and separate bootm steps. FIT allows U‑Boot to load and verify one image and reduces scripting overhead and redundant memory copies. Yocto and U‑Boot tooling support producing FIT images with kernel+DTB+initramfs. Example U‑Boot snippet (production env): setenv bootdelay 0 setenv autoload n setenv bootcmd 'fatload mmc 0:1 ${kernel_addr_r} zImage; fatload mmc 0:1 ${fdt_addr_r} devicetree.dtb; booti ${kernel_addr_r} - ${fdt_addr_r}' saveenv Reference: U‑Boot environment and SPL guidance. Make the kernel and initramfs faster: compression, initcalls and modules This is where you trade size, memory and CPU for latency. Two heavy hitters are kernel decompression and module/driver initialization. Compression tradeoffs - Modern kernels support several compression formats. Recent work added zstd support to kernel/initramfs; zstd typically yields better decompression speed than xz and better size thangzip , whilelz4 often yields the fastest decompression but at worse ratio. The kernel patches and community testing (including large deployments) show zstd as an attractive sweet spot; in real deployments Facebook reported large reductions in initramfs decompression time when switching to zstd. - Practical rule: test on your target SoC. On low‑power devices the decompressor speed and cache configuration matter; on fast application processors the size reduction (improving cache/memory footprint) can also beat raw decompression time. Compression snapshot (representative, taken from kernel discussion and test reports): | Algorithm | Typical compressed kernel size (x86_64 example) | Decompression notes | |---|---:|---| | none (uncompressed) | 32.6 MB | No decompression cost but larger RAM/copy time | | lz4 | 10.7 MB | Very fast decompress; tradeoff: larger than zstd | | zstd | 7.4 MB | Good ratio and very fast; often best overall tradeoff | | gzip | 8.5 MB | Moderate speed and ratio | | xz / lzma | 6.5-6.8 MB | Best ratio in many cases, slowest decompress | Kernel initcalls and module strategy - Use initcall_debug during profiling, find the top initcalls by duration, and decide whether to:- Move slow, non‑critical init work to later (defer via late_initcall or userspace), - Build it as a module and load from a minimal initramfs or userspace script, or - Keep it builtin if filesystem access delays would otherwise hold the system up. - The trade is not binary: moving a driver to modules removes its initcall from kernel boot, but module loading can still block userspace and hit slow storage or udev. Measure both kernel and userspace timelines before changing strategy. Initramfs slimming and bundling - Make the initramfs as tiny as practical: a busybox -based init with only the scripts and device nodes needed to mount the real root (or to start the minimal services you want available at that point). Buildroot and Yocto have features to produce tiny initramfs images and to bundle them into FIT images. Embedding initramfs into the kernel avoids a separate ramdisk load step (it becomes part of the kernel image load/unpack). - When using compressed root filesystems, pick the one that fits your device constraints: a read‑only compressed squashfs for immutable systems,UBIFS for writable raw NAND with fast mount (UBIFS avoids full media scan and mounts far faster than JFFS2), orext4 on eMMC with tuned mount options. Practical knobs to try (example kernel command line for testing profiling): console=ttyS0,115200 earlyprintk=serial,ttyS0,115200 printk.time=1 initcall_debug loglevel=3 Trace, decode with dmesg | grep initcall and act on the top offenders. Service ordering and filesystem tricks that shave seconds Userspace ordering and filesystem mounting are often the last visible stretch before shell. Service parallelization - Let the init system run services in parallel and use activation primitives: - With systemd , rely on socket activation and correct unitType= values (Type=notify ,Type=dbus ,Type=forking where appropriate) so systemd can parallelize work and not wait unnecessarily. Socket‑based activation lets services appear available while they start in the background. Usesystemd-analyze to find expensive, blocking units. - Avoid blanket network-online.target waits unless the product explicitly requires network at boot. Many services block on the network because ofNetworkManager-wait-online orifup@.service . Replace waiting with on‑demand approaches or a short timeout. - With - Use systemd-analyze blame andcritical-chain to identify the dependency chain that actually determines your time‑to‑shell. Often a single service waiting ondbus or DHCP accounts for most of the delay. Filesystem and driver tricks - Mount options: disable atime bookkeeping (noatime ), considerdata=writeback only when acceptable, and tunecommit= to reduce sync pressure for boot‑critical partitions. These reduce writes and metadata pressure early in boot but carry durability tradeoffs. Check the mount man page for exact semantics.
Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.