Comment by gojkoa

Back to stories Open in a window

Comment by gojkoa

I'm working for myself (own product, decently profitable SaaS, been in production for 7 years now and used by actual customers), so I'm not sure if I fit your "employed" category or not. If I do, then here's my setup:

1. I have a custom-built minimal flow of 5 "commands", that take an idea to implementation through capturing intent, impact analysis, iterative planning, execution orchestration and then review/demo. each of these effectively builds up a single shared plan file. The intent command takes one or two lines of free text and turns it into a structured intent document (goals/anti-goals/constraints...); impact-analysis takes the intent document and maps it to the codebase, looking for functional gaps and things that need to change. iterative-planning takes the result of the impact analysis and splits into tasks/phases that are independently verifiable and deployable, and builds a task list... so they all build on each other, and update the same plan file that sits in git and I review it as it goes through the pipeline

2. we have a minimal CONTRIBUTING.md that explains the shape of the workspace and the key rules how to work, that's applicable to humans and agents. CLAUDE.md loads it from @CONTRIBUTING.md

3. the guardrails of what the agent is allowed or not allowed to do mostly sit in deterministic tools, such as custom linting rules, custom style checks, and they are all invoked from eslint via the custom language plugin. this has grown to thousands of rules, linting programming language code but also html, scss, yaml, liquid, markdown.... Eslint runs as a post-tool use hook on edit and write, so each file an agent writes gets immediate feedback and fixes. when we catch the agent doing something it should have not done with the code or docs, we get it to write another custom rule with unit tests for the rule. With each rejection Claude also gets a helpful message what to do instead.

4. there's a "regulator" script that helps us avoid decision fatigue for approvals. it runs as a pre-tool use hook for Bash commands, does deep parsing of whatever sausage Claude wants to run, goes into loops, function definitions etc, then goes through our rule set and approves/rejects or forces an ask. With each rejection Claude also gets a helpful message what to do instead. (e.g. don't run npx, use eslint directly from the path). Each time claude asks, we run the command through a debug script to understand why it's asking, and add another rule.

5. the latest addition is an orchestrator script that takes our plan file format and turns into Claude Dynamic Workflow descriptions deterministically, and parses the progress of the workflows so suggest what's slow and what can be moved out of LLM processing to deterministic tools. This significantly reduced the token spend (now running at about ~20% of the token spend for workflows before) and time (from average 5-6 hours per workflow to about 20-30 minutes). It also removed the 10-15 minute wait that we had while claude was LLM constructing workflows from our plans.

6. there's a demo command that flies through the user interface based on the plan to demonstrate what users see with the new version, recording it to webm using playwright, and I can play it at a higher speed to quickly get an overview what an agent did.

7. we tend to look for ways to get faster feedback on things that agents repeatedly do badly, such as UX or UI changes. We have a set of static HTML pages with the visual design language, showing styling for key elements and components, and a set of static demo pages showing key application pages with realistic data in lots of different states. as part of the impact analysis, agents will update demo pages or add new ones so we can review/complain. there's a script that audits demo pages for WCAG and other styling issues. another example are end-to-end api tests, which evolved massively to prove api contracts but also allow agents to get their own feedback and troubleshoot quickly. generally, divide and conquer for feedback.

This tends to work generally well. I feel productive. I still review most code when it completes via git diff, but it's mostly clean because the linting rules are forcing it to write code the way I want it to be written. We use a method based on the attribute-component-capability matrix to figure out what needs manual exploratory testing and how much, and do that in addition to automated tests when needed.

happy to provide any more info if you're interested.

Replies gojkoa · 26d
Open on HN
Loading the discussion…

Domain filters

Stories from these domains are hidden from every list. Subdomains match too: blocking substack.com also hides danluu.substack.com.

    About YAVCHN

    YAVCHN is a reader for Hacker News and Lobsters, with articles and discussions in separate windows or Classic pages.

    Created by Paul Parks and built with PUDL.

    YAVCHN source code on GitHub

    Help

    Keyboard

    j / k
    Move down and up the story list. The arrow keys scroll whatever has focus.
    Enter
    Read the marked story in the article reader.
    ]
    Read the next story in the same article-reader applet. Back returns to the previous story.
    p
    Pin or unpin the marked story, which keeps it in Pinned.
    n / N
    Move to the next or previous top-level comment in the window in front.
    c
    Collapse or expand that comment.
    f
    Hide or show the story list.
    Esc
    Close a menu or this help.
    Access key m
    Go to the menu bar. Most browsers take it with Alt on Windows and Linux, and Safari with Control and Option.
    ?
    Show this help.

    Windows

    Each story opens in a window holding its article above its discussion; drag the bar between them to share the room differently. A window can be moved by its title bar, resized from any edge, snapped to a half or a corner by dragging it there, maximised, or minimised to the bar at the foot of the page. Use Window > New reader window to open an empty reader, or Story > Open in new reader window to open another reader for the current article. Docked readers keep their articles when you select another story from the sidebar. Minimized readers can be restored and reused for their site. A window's Next story link reads on down the list in the same window.

    A link in a comment or an article to another Hacker News or Lobsters thread opens that thread in a window too. A link to a single HN comment opens the comment above its replies.

    While a story's window is in front, the Story and Discussion menus in the menu bar hold its commands: pinning, Next story, sorting, collapsing every thread, jumping to the first new comment. Each window also remembers where you were in its article and discussion, so a reload, or Back to a story that Next took you past, finds your place again. Closing a window forgets it.

    The whole arrangement lives in the address, so a bookmark or a shared link brings it back, and Back undoes the last change. Moving between Hacker News, Lobsters, their lists, Pinned and Find changes only the list, and leaves the windows open.

    The list

    The pin at the start of a row keeps the story in Pinned, and the cross at its end hides it. Pinned can be narrowed by words in the title, site or author, by source, and to the stories you haven't opened yet, and ordered by when you pinned them, by points or by comments; the filters are part of the address, so a filtered view can be bookmarked. Scroll past the end of the list to load more. Domain filters, in the View menu, hide every story from a site.

    Browsing view

    View > Windowed and View > Classic select the browsing view and save your default in this browser. Window view reuses a reader for each feed. Classic view opens stories and applets as pages. Open as a page and Open in a window are one-off actions that do not change your saved default. Direct page links always open as pages.

    Applets

    The Applets menu in the menu bar holds three tools, each a window of its own. Replies to me takes your Hacker News user name and lists the replies to your last thirty comments and stories, checking again every three minutes while it is open, and marking what is new since you last marked them read. Look up a user opens a profile on Hacker News or Lobsters, with their submissions and recent comments, as a commenter's name in any discussion does; the bar at the top of a profile looks up someone else in the same window, and Back returns to the one before. Who is hiring? filters the posts of HN's monthly hiring threads by the words you type.

    They read only what the sites publish to everyone, so none of them asks for a login, and your user name stays in this browser.

    Find

    Find takes any link and lists every time it was submitted to Hacker News and Lobsters, so you can read each discussion of it.

    About

    YAVCHN never sees your Hacker News or Lobsters login. The discussion is fetched from each site's public API; to vote or reply, follow the link above the discussion, or the arrow beside a comment, to the source's own site. Pins, hidden stories, filters and layout are kept in this browser only.

    Open source: github.com/paulmooreparks/yavchn. Built with PUDL.