Show HN: I spent $450 on GCP's Video API, so I built a local alternative

Show HN: I spent $450 on GCP's Video API, so I built a local alternative

Hey!

I have 2TB+ of personal video footage. Finding specific moments was impossible, like searching PDFs by content, but for video. Google's Video Intelligence API worked, but it cost me $450+ for just a few videos, and I'd have to upload all my personal footage to their cloud.

So I built Edit Mind to do it locally.

The core problem:

You can search PDFs by their content in Finder (Mac OS). Why can't we do the same with videos? Every cloud solution either costs a fortune at scale or requires uploading your personal footage to cloud storage but I don’t wanna have my personal raw videos uploaded to the cloud.

How it works:

- Everything runs on your machine (your raw videos never leave your computer) - Indexes videos once: transcribes audio, detects objects (YOLO), recognizes faces, analyzes emotions - Stores metadata in ChromaDB (local vector database) - Natural language gets parsed into structured queries, then semantic search finds matches (uses Gemini API currently, but can swap in a local LLM like Ollama if you prefer)

What it actually does:

1. Type: "scenes where @Ilias is looking happy, eating a pizza" 2. Behind the scenes it converts this to: {"faces":["Ilias"], "emotions": ["happy"], "objects": ["pizza"]} 3. Then searches your local vector database for matching scenes across 2TB of footage in seconds.

Real cost comparison:

* GCP Video Intelligence API: ~$0.10/minute of video (https://cloud.google.com/video-intelligence/pricing) = $100+ for 200 (5 minutes) videos (I have over 3000 videos), this will cost me over 1500+$ for getting the videos analyzed * Edit Mind: Free after initial setup, runs on your own hardware (you need to have good one to handle this heavy video processing tasks)

Technical choices:

Built with Electron because I needed real filesystem access and didn't want browser storage limits. Python backend handles the heavy ML work (face_recognition, YOLOv8, FER for emotions) and communicate via web sockets. The analysis pipeline is plugin-based – took me less than a day to add color dominate as a separate plugin.

Current limitations:

* Needs decent hardware (GPU recommended but not required) * Face recognition requires some manual training (adding known faces) * UI is functional but could be prettier * Query parsing uses Gemini API by default (but you can configure it to use local alternatives)

Why I'm sharing this:

I can't be the only person with this problem. Videographers, parents with years of family footage, documentary filmmakers, anyone with large video libraries. The code is MIT licensed and designed to be extended.

Would love feedback on:

1. Is $450+ for cloud video analysis a common pain point, or am I an outlier? 2. What other analyzers would be useful? (thinking about camera movement analysis and scene type classification like POV/vlog/interview) 3. Should I prioritize making it 100% offline by default, or is the Gemini API option fine for most use cases?

GitHub: https://github.com/iliashad/edit-mind

Demo: https://youtu.be/Ky9v85Mk6aY

Built this over a few weekends because I was tired of paying Google to search my own videos. Happy to discuss architecture decisions or the ML pipeline!

Discussion 0 comments · 4 points · iliashad · 2025-10-24
Open on HN
Loading the discussion…

Domain filters

Stories from these domains are hidden from every list. Subdomains match too: blocking substack.com also hides danluu.substack.com.

    About YAVCHN

    YAVCHN is a reader for Hacker News and Lobsters, with articles and discussions in separate windows or Classic pages.

    Created by Paul Parks and built with PUDL.

    YAVCHN source code on GitHub

    Help

    Keyboard

    j / k
    Move down and up the story list. The arrow keys scroll whatever has focus.
    Enter
    Read the marked story in the article reader.
    ]
    Read the next story in the same article-reader applet. Back returns to the previous story.
    p
    Pin or unpin the marked story, which keeps it in Pinned.
    n / N
    Move to the next or previous top-level comment in the window in front.
    c
    Collapse or expand that comment.
    f
    Hide or show the story list.
    Esc
    Close a menu or this help.
    Access key m
    Go to the menu bar. Most browsers take it with Alt on Windows and Linux, and Safari with Control and Option.
    ?
    Show this help.

    Windows

    Each story opens in a window holding its article above its discussion; drag the bar between them to share the room differently. A window can be moved by its title bar, resized from any edge, snapped to a half or a corner by dragging it there, maximised, or minimised to the bar at the foot of the page. Use Window > New reader window to open an empty reader, or Story > Open in new reader window to open another reader for the current article. Docked readers keep their articles when you select another story from the sidebar. Minimized readers can be restored and reused for their site. A window's Next story link reads on down the list in the same window.

    A link in a comment or an article to another Hacker News or Lobsters thread opens that thread in a window too. A link to a single HN comment opens the comment above its replies.

    While a story's window is in front, the Story and Discussion menus in the menu bar hold its commands: pinning, Next story, sorting, collapsing every thread, jumping to the first new comment. Each window also remembers where you were in its article and discussion, so a reload, or Back to a story that Next took you past, finds your place again. Closing a window forgets it.

    The whole arrangement lives in the address, so a bookmark or a shared link brings it back, and Back undoes the last change. Moving between Hacker News, Lobsters, their lists, Pinned and Find changes only the list, and leaves the windows open.

    The list

    The pin at the start of a row keeps the story in Pinned, and the cross at its end hides it. Pinned can be narrowed by words in the title, site or author, by source, and to the stories you haven't opened yet, and ordered by when you pinned them, by points or by comments; the filters are part of the address, so a filtered view can be bookmarked. Scroll past the end of the list to load more. Domain filters, in the View menu, hide every story from a site.

    Browsing view

    View > Windowed and View > Classic select the browsing view and save your default in this browser. Window view reuses a reader for each feed. Classic view opens stories and applets as pages. Open as a page is a one-off action that does not change your saved default. Use the Windowed selector to return an article to a window. Direct page links always open as pages.

    Applets

    The Applets menu in the menu bar holds three tools, each a window of its own. Replies to me takes your Hacker News user name and lists the replies to your last thirty comments and stories, checking again every three minutes while it is open, and marking what is new since you last marked them read. Look up a user opens a profile on Hacker News or Lobsters, with their submissions and recent comments, as a commenter's name in any discussion does; the bar at the top of a profile looks up someone else in the same window, and Back returns to the one before. Who is hiring? filters the posts of HN's monthly hiring threads by the words you type.

    They read only what the sites publish to everyone, so none of them asks for a login, and your user name stays in this browser.

    Find

    Find takes any link and lists every time it was submitted to Hacker News and Lobsters, so you can read each discussion of it.

    About

    YAVCHN never sees your Hacker News or Lobsters login. The discussion is fetched from each site's public API; to vote or reply, follow the link above the discussion, or the arrow beside a comment, to the source's own site. Pins, hidden stories, filters and layout are kept in this browser only.

    Open source: github.com/paulmooreparks/yavchn. Built with PUDL.