FolderPilot: A Local AI Folder Organizer That Is Built Never to Delete a File
DEV Community

FolderPilot: A Local AI Folder Organizer That Is Built Never to Delete a File

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built FolderPilot is a local web app that scans any folder, reads what is inside the files, and proposes a cleaner structure. It shows the result as an interactive, color-coded map. You review the plan, edit it, and approve it. Nothing changes on disk until you do. I built it for a friend whose folders hold exactly the kind of files you do not want to upload anywhere: ID scans, marksheets, resumes, offer letters, and assignments, all mixed together with installers and screenshots. Cloud organizers want those documents on someone else's server. Cleanup scripts delete things. I wanted a tool that does neither. Three rules shaped the whole project: - Everything runs on 127.0.0.1 with an open-weight model served by Ollama. - The code has no delete, unlink, remove, or truncate path, so there is nothing to bypass. - Every change is journaled before it happens, so any operation, batch, or the whole workspace can be undone. Demo The video walks through scanning a real folder, exploring the treemap and sunburst views, reviewing the before and after plan, applying it, and rolling it back. It also shows the folder chat refusing a request to delete files. Code TejasRawool186 / FolderPilot A local, privacy-first web app that analyzes any folder, shows it visually with color-coded suggested sorting, lets you customize and approve the plan, and lets you chat with the folder. It can never delete anything. FolderPilot Organize anything. Delete nothing. Privacy-first, local folder organizer with interactive visual analytics, append-only rollback journaling, and on-device AI. Overview · How it works · Features · Safety model · Quick start Why this exists Personal directories such as Downloads, Desktop, and work folders routinely accumulate sensitive files-including tax returns, identity cards, college transcripts, and resumes. Most of these files remain disorganized because manual sorting is slow and tedious. Typical cloud organizers require uploading private documents to external servers. Other utility scripts use aggressive delete operations that risk permanent data loss. FolderPilot runs entirely on local hardware, inspects documents privately on 127.0.0.1 , and guarantees that no file can ever be deleted. Why open-source, local AI Running open-weight language models locally on consumer hardware changes the economics and security of desktop file management: | Dimension | Local Open-Source AI (FolderPilot) | Cloud AI APIs | |---|---|---| | Privacy | Local-only. File contents, extracted text, and metadata never | The repository has a FastAPI backend, a Next.js frontend, a pytest suite for the safety guarantees, and docs covering the architecture, the AI integration, and the API. The Problem A Downloads folder is where files go to pile up. After a year it holds duplicate resumes (resume_final , resume_final2 ), screenshots named IMG_2938.png , installers you already ran, and a few documents you would be uncomfortable seeing leaked. Existing options each fail in a different way: - Cloud organizers need your documents uploaded, including identity cards and transcripts. - Cleanup scripts sort by extension, which cannot tell an assignment from a bank statement, and some delete aggressively. - Doing it by hand is slow enough that nobody finishes. The Idea Make the AI the last resort, and make the file system layer safe by construction. Most files do not need a language model. An extension, a hash, or a keyword count settles them. So FolderPilot runs cheap checks first and only sends the genuinely ambiguous files to a small local model. That keeps it usable on a laptop with 8 GB of RAM and no GPU, which is the machine I developed it on. On the safety side, I did not add a "confirm delete" dialog. I removed deletion from the program. The file operations module only exposes mkdir , move , and rename . How It Works The scanner walks the chosen folder and streams progress to the UI. Each file then goes through four tiers, stopping at the first one that is confident. flowchart TD A["Select any folder"] --> B["Background scan"] B --> C["Tier 1: extension and filename rules"] C --> D["Tier 2: SHA-256 duplicate detection"] D --> E["Tier 3: keyword prototypes on extracted text"] E --> F["Tier 4: local LLM, JSON schema output"] F --> G["Draft plan, unapproved"] G --> H["User reviews, edits, approves"] H --> I["Journal entry written"] I --> J["mkdir / move / rename"] J --> K["Undo: single op, batch, or full workspace"] | Tier | Mechanism | Handles | |---|---|---| | 1. Rules | Extension taxonomy and filename regex | Code, archives, media, installers, screenshots | | 2. Hashes | Lazy SHA-256 on files with equal sizes | Exact duplicates and repeated downloads | | 3. Prototypes | Keyword frequency in the first ~500 extracted tokens | Assignments, invoices, offer letters | | 4. Local LLM | Ollama with constrained JSON output | Ambiguous PDFs and poorly named files | Every file gets a category and a reason, so the review screen can show why it was placed where it was. Technical Architecture | Layer | Technology | Why | |---|---|---| | Frontend | Next.js 16, React 19, Tailwind, D3.js | D3 gives full control over the treemap, sunburst, and tree comparison | | Backend | Python 3.11, FastAPI | Async routes and server-sent events for live scan progress | | Storage | SQLite in WAL mode | File index, plans, and the journal in one local file | | Extraction | PyMuPDF, python-docx, python-pptx | Reads text from PDFs, Word, and PowerPoint files | | Local AI | Ollama with Gemma 3 1B (default), Qwen 2.5 1.5B | Small enough for CPU-only machines | The frontend never talks to anything but the local backend, which binds to 127.0.0.1 . The interface has these views: - Treemap with drill-down zoom and a spotlight filter. - Sunburst view of size by depth. - Before and after tree showing the proposed structure next to the current one. - Duplicate clusters with keep-original and keep-newest suggestions. - Chaos score from 0 to 100, based on root-level clutter, nesting imbalance, and duplicate ratio, with a slider that simulates cleanup. - In-app previewer for images, audio, video, PDFs, Office documents, spreadsheets, and a hex view for binaries. How I Used Open-Source AI The default model is Gemma 3 1B (gemma3:1b , Q4_K_M, roughly 1.2 GB of RAM), served through Ollama. Qwen 2.5 1.5B is supported as an alternative. The model can be switched from the navigation bar at runtime, through the POST /api/ai/model endpoint, or with the FOLDERPILOT_OLLAMA_LLM_MODEL environment variable. The model does two jobs: - Fallback classification. For files the first three tiers cannot place, it returns JSON matching a fixed schema. Constraining the output keeps replies short and parseable. - Folder chat and summaries. Counts and size questions are answered from SQLite, and the model handles wording and short document summaries. Why open weights mattered here: | Local open model | Cloud API | | |---|---|---| | Privacy | File text and metadata stay on the machine | Excerpts of private documents leave the machine | | Cost | No per-file cost | Metered by volume | | Offline | Works once weights are cached | Needs a connection | | Model choice | Swap models in one setting | Tied to one provider | The tradeoff is accuracy. A 1B model will misclassify some ambiguous files, especially ones with little extractable text. That is why low-confidence files are surfaced in a review queue and why nothing is applied without approval. A user's manual category overrides are stored in local SQLite to guide later scans. Real Example Example 1: classification A file named assignment_final.pdf sits next to assignment.pdf . The extension alone says "document". Tier 2 notices the sizes differ, so they are not exact duplicates. Tier 3 reads the first page, finds terms common to coursework, and places both under a college category. The plan shows the reason on the card, and the user can drag them elsewhere if it is wrong. Example 2: a destructive request In the folder chat, I typed a request to delete all the duplicate files. The chat router intercepts deletion and removal requests before they reach the model. The reply explains that FolderPilot never deletes, and offers to move the duplicates into a _Review_Later/ folder instead. A test (test_chat_delete_refusal ) covers this behavior. Example 3: undo After applying a plan, the Journal view lists every operation. One click on a single operation restores that file. One click on the batch restores everything the plan touched. Because the journal entry is written before each operation, an interrupted run leaves a consistent record. What Makes It Different Most organizers confirm before they delete. This one cannot delete. - Safety by absence. safe_ops.py has no delete, unlink, remove, or truncate path.test_no_delete_functions_exist checks that this stays true. - No overwrites. Name collisions get a suffix, as in Document (1).pdf . - Journal before action. The undo log is not an afterthought. Entries exist before the file system changes. - Protected paths. Folder browsing rejects system locations such as C:\Windows andC:\Program Files , and has tests for path traversal and symlink escapes. - LLM as the last tier. The model sees only the files the cheaper checks could not settle, which is what makes a 1B model enough. Build Process I started by writing a requirements document before any code, because the safety rules had to be requirements and not afterthoughts. The repo's docs/ folder keeps that specification along with architecture notes and a development log. The backend came first: scanner, safe file operations, journal, and undo. I built the visual layer after that, once the plan and journal data was reliable. The default model started as Qwen 2.5 1.5B and moved to Gemma 3 1B after I tuned for memory use and speed on an 8 GB machine, with Qwen kept as a switchable option. The t

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.