I implemented the YOLO26n model inference from scratch using ARM64 Assembly Language (No framework) [P]
Implementation Overview
This was my Bachelor's Final Project: implementing YOLO26n inference completely from scratch using ARM64 Assembly Language and C, without relying on existing inference frameworks. The goal was to understand how modern neural network inference engines work at a low level and explore optimization techniques for faster and more efficient edge AI execution on Raspberry Pi 4.
Implementation Includes
- ARM64 Assembly Language + C inference engine
- ARM NEON SIMD optimization
- Winograd convolution
- Optimized GEMM kernels
- Cache-aware tiling
- Custom ARM64 micro-kernels
- Operator fusion
- Attention mechanism
- YOLO26 components: Conv, C3K2, SPPF, C2PSA, PSA, BottleNeck, and Detect
Memory Layout and Model Extraction
I extracted the YOLO26n model parameters and redesigned the memory layout into a custom binary format optimized for the inference pipeline.
Results and Request for Feedback
The implementation produces correct object detection results, but the performance improvement was lower than I initially expected. I would appreciate feedback and suggestions from anyone about:
- CNN inference optimization
- ARM NEON/vectorization
- Memory layout and cache optimization
- Low-level neural network acceleration
Repository
Repository: https://github.com/mohammad-ghaderi/YOLO26
Thanks for any feedback or suggestions.
Comments
No comments yet. Start the discussion.