DotGhostBoard 2.1 'Leviathan' - Smart Clipboard Tagging & Contextual Actions on Linux (Part 1)
DEV Community

DotGhostBoard 2.1 'Leviathan' - Smart Clipboard Tagging & Contextual Actions on Linux (Part 1)

How we built a rules-based clipboard intelligence engine for Linux - local, deterministic, sub-millisecond for typical payloads (< 5KB), zero AI, zero telemetry. Auto-tagging, contextual smart actions, and a tight pipeline architecture explained.

You copy a minified JSON blob from a terminal log. To actually read it, you paste it into an editor, run a formatter, then copy the result back. You copy an API key from a config file and your clipboard manager stores it as plain text, alongside screenshots and grocery lists, until you manually hunt it down and secure it.

Most clipboard tools treat every copied item the same: inert text in a database. For anyone who spends time in terminals, editors, and browser devtools, that uniformity creates constant small frictions that accumulate fast.

DotGhostBoard 2.1 "Leviathan" takes a different approach. The clipboard understands what it's holding and surfaces the right action immediately - format this JSON, open that URL, secure this credential.

This is a two-part series on how we built it:

  • Part 1 (you're here): Rules-based Smart Auto-Tagging, Contextual Smart Actions, and the pipeline behind them.
  • Part 2: Cryptographic .vault Backups, CSPRNG Password Generator, Secret Expiry, and Encrypted Version History.

The Design Constraint: Local, Deterministic, Fast

The obvious shortcut for text classification is an LLM or a cloud API. For a clipboard manager, that's a non-starter on three counts:

  1. Trust boundary. Your clipboard routinely holds API keys, internal paths, private tokens, and proprietary code snippets. Sending that to any remote service - even a local on-device model - introduces a trust boundary that shouldn't exist in a tool people actually rely on for security.
  2. Memory. Loading inference libraries adds hundreds of MB of idle RAM.
  3. Latency. Any latency between Ctrl+C and the item appearing in history is enough to make the tool feel broken.

Our constraint was concrete: every classification must run locally, produce deterministic results, and complete in under 1ms for typical clipboard payloads (< 5KB), scaling to ~1.4ms worst-case at the 8KB classification cap.

The approach: pre-compiled regular expressions, structural heuristics, and a pool-based keyspace entropy estimate - no opaque models.

1. The Auto-Tagger: How Classification Actually Works

The engine lives in core/security/auto_tagger.py. It runs synchronously on every new clipboard item before storage and UI rendering.

Captured Text ──โ–บ Pre-compiled Patterns & Heuristics ──โ–บ Tags ──โ–บ Storage & UI
├── URL regex ──โ–บ #link
├── JSON fast-path + json.loads ──โ–บ #json
├── Code keywords + bracket density ──โ–บ #code
├── Known prefixes + entropy gate ──โ–บ #secret
├── IPv4 regex + IPv6 regex ──โ–บ #ip
├── RFC-like email regex ──โ–บ #email
├── Unix / Windows path regex ──โ–บ #path
└── Hex length check {32,64} ──โ–บ #hash

All patterns compile once at module load time. The function is pure - no I/O, no Qt, no database access - which makes it fast to call and trivial to test.

Here's the real implementation, abbreviated:

# core/security/auto_tagger.py
import json, re, math, string

# ── Compiled patterns (module-level - compiled once, reused forever) ──────────
_RE_URL = re.compile(r"https?://[^\s\"'<>]{3,}|ftp://[^\s\"'<>]{3,}", re.IGNORECASE)
_RE_EMAIL = re.compile(r"\b[A-Za-z0-9._%+\-]+@[A-Za-z0-9.\-]+\.[A-Za-z]{2,}\b")
_RE_IPV4 = re.compile(  # validates 0-255 per octet
    r"\b(?:(?:25[0-5]|2[0-4]\d|[01]?\d\d?)\.){3}"
    r"(?:25[0-5]|2[0-4]\d|[01]?\d\d?)\b"
)
_RE_IPV6 = re.compile(r"\b(?:[0-9a-fA-F]{1,4}:){2,7}[0-9a-fA-F]{1,4}\b")

# Pragmatic range matching MD5 (32), SHA-1 (40), SHA-256 (64); strict lengths planned for v2.2
_RE_HEX_HASH = re.compile(r"\b[0-9a-fA-F]{32,64}\b")

# Known credential prefixes - the first line of secret detection
_RE_SECRET_PATTERNS = [
    re.compile(r"ghp_[A-Za-z0-9]{36,}"),           # GitHub PAT
    re.compile(r"AKIA[0-9A-Z]{16}"),               # AWS Access Key
    re.compile(r"sk-[A-Za-z0-9]{32,}"),            # OpenAI / Stripe
    re.compile(r"xox[baprs]-[A-Za-z0-9\-]{10,}"),  # Slack tokens
    re.compile(r"eyJ[A-Za-z0-9\-_]{10,}\.[A-Za-z0-9\-_]{10,}"),  # JWT
    re.compile(r"-----BEGIN (?:RSA |EC |OPENSSH )?PRIVATE KEY-----"),
]

# (Note: _RE_PATH, _CODE_KEYWORDS, and _CODE_BRACKETS are omitted here for brevity)

# ── Pool-based keyspace entropy estimate ──────────────────────────────────────
def _entropy_bits(text: str) -> float:
    """Estimate theoretical keyspace entropy: H = len * logโ‚‚(pool_size)."""
    pool = 0
    if any(c in string.ascii_lowercase for c in text): pool += 26
    if any(c in string.ascii_uppercase for c in text): pool += 26
    if any(c in string.digits for c in text): pool += 10
    # Any non-alphanumeric character (symbols, punctuation, or non-ASCII) adds 32
    if any(c not in string.ascii_letters + string.digits for c in text): pool += 32
    return math.log2(pool or 26) * len(text)

# ── Main classification function ──────────────────────────────────────────────
def detect_tags(content: str) -> list[str]:
    if not content or not isinstance(content, str):
        return []
    tags: list[str] = []
    text = content.strip()

    if _RE_URL.search(text): tags.append("#link")
    if _RE_EMAIL.search(text): tags.append("#email")
    if _RE_IPV4.search(text) or _RE_IPV6.search(text): tags.append("#ip")
    if _RE_PATH.search(text) and "#link" not in tags: tags.append("#path")
    if _RE_HEX_HASH.search(text): tags.append("#hash")

    # Secret detection: known prefix first, entropy gate as fallback
    is_secret = any(p.search(text) for p in _RE_SECRET_PATTERNS)
    if not is_secret and len(text) >= 16 and " " not in text and "\n" not in text:
        if _entropy_bits(text) >= 90:
            is_secret = True
    if is_secret: tags.append("#secret")

    # JSON: only attempt parse if it looks like an object or array
    if text.lstrip().startswith(("{", "[")):
        try:
            json.loads(text)
            tags.append("#json")
        except (json.JSONDecodeError, ValueError):
            pass

    # Code: multi-line with language keywords or dense bracket patterns
    if "\n" in text and len(text) > 40:
        if _CODE_KEYWORDS.search(text) or _CODE_BRACKETS.search(text):
            tags.append("#code")

    return list(dict.fromkeys(tags))  # deduplicate, preserve order

Secret Detection: Two Layers

The #secret classifier uses a two-step approach:

  • Layer 1 - Known prefixes. Patterns like ghp_, AKIA, eyJ (JWT header) are unambiguous. Match any of them and the tag is assigned immediately, no entropy calculation needed.
  • Layer 2 - Entropy gate. For everything else: single-line, no spaces, at least 16 characters, and a pool-based keyspace entropy ≥ 90 bits.

The 90-bit threshold was tuned empirically to catch random high-entropy tokens, API keys, and strong bearer credentials while skipping short identifiers, dictionary words, and common camelCase variable names.

Both conditions guard against a common pitfall: pure Shannon entropy on raw text produces too many false positives on long sentences or dense code snippets. The entropy function here is H = len × logโ‚‚(pool_size) - a measure of the theoretical keyspace, not character frequency distribution.

Note on non-ASCII characters: because any(c not in string.ascii_letters + string.digits ...) catches non-ASCII characters (e.g. Arabic script, accented letters, emoji), unicode text will trigger the +32 symbols branch. This heuristic is tuned primarily for ASCII credentials; multilingual script entropy is handled gracefully without throwing exceptions.

Known Limitations

Honest engineering requires talking about edge cases:

  • Long paths tagged as #secret: A Unix path like /home/kareem/StudioProjects/DotGhostBoard/main.py currently receives both #path and #secret - it satisfies the entropy gate because it has no spaces and mixes uppercase, lowercase, digits, and /. We've logged this as a known limitation; v2.2 will add #path as an explicit exclusion condition for the entropy fallback.
  • Hashes flagged as #secret: A 32-character hex string (like an MD5 digest) uses digits + lowercase (pool = 36), yielding 32 × logโ‚‚(36) ≈ 165 bits, which clears the ≥ 90-bit gate. Consequently, MD5 and SHA digests receive both #hash and #secret; a dedicated check to prioritize #hash alone is slated for v2.2.

2. What the Tag Enables: Contextual Smart Actions

Tags are displayed as colored pill chips on each card. More importantly, they drive Contextual Smart Actions - buttons that appear dynamically in the card toolbar based on what was detected.

Tag Action What it does
#link ๐Ÿ”— Open Link QDesktopServices.openUrl() - no copy-paste to browser

| #json | {

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.