Comment by chronicallyoffl

Back to stories Open in a window
Comment

Comment by chronicallyoffl

hi folks,

i'm an engineer and amateur hacker who's been working on a little side project, and i want to show it off and also solicit your feedback.

i'm building efficient machine learning models for natural language processing on consumer cpus: sentiment analysis, task routing, document retrieval, etc. i know there are a lot of models for these tasks already, but too many of them are:

(1) difficult to configure, or (2) parameter-inefficient, or (3) too high-latency for most applications.

instead of building academically "novel" architectures, my project is making parameter-efficient and low-latency models easy to use through distillation or finetuning, computation graph optimization, and quantization. the eventual vision is a single terminal command to install the model. (i've not quite achieved this, it's like three or four commands right now).

i'm posting here because i've completed a beta version of my first model and i'd be very grateful for community feedback. the first model is called `offline-sentiment-small` and it's for binary sentiment analysis (positive/negative). you can install the runtime using `pip install offlinedemo` and download the checkpoint from the releases page of

https://github.com/offlineisbetter/offlinedemo

the readme of the above repository shows benchmarks against three other models for sentiment analysis: distilbert, roberta, and modernbert. they all do about the same on sst2 (~0.92-0.95 f1 score) because it's a relatively easy dataset. you'll notice that, though my model is 230m parameters, you need to go down to distilbert (67m) to get the same kind of latency. i want to stress-test these models on datasets with longer-context or more difficult sentiment tasks, to see how performance degrades with both my model and these models. if you know about any good datasets for this, please do let me know!

for those curious, my model is a liquid foundation model (lfm2.5) encoder (on hf, LiquidAI/LFM2.5-Encoder-230M) finetuned by low-rank adaptation on the sst2 training set, optimized with onnx, and quantized to int8. i chose lfm over bert as the base model because it uses grouped-query attention and convolution layers instead of dense attention. these changes over the "vanilla" transformer architecture make its latency subquadratic in the input context (and i've verified this empirically). lfms are often considered more parameter-efficient than transformers: for example, liquid ai's 2.6b autoregressive model competes with llama 7b.

and disclaimer, i have no affiliation to liquid, i just think their work is cool!

i want to emphasize that i'm not claiming technical novelty! rather, i'm building a way to make parameter-efficient, distilled/finetuned, and quantized models more accessible to more people. right now, if you wanted to run a distilled and quantized text model, it takes a non-negligible amount of compute and time and effort to set it up. not everybody wants to do that, and so out of convenience they'll just opt for a big llm in the cloud.

i worry that so many people these days pay for openai/anthropic tokens just to do a simple task, such as sort their emails into categories. not only is this wasteful, high-latency, bad for the environment, etc. but you're giving somebody else your data. so i've started the offlineisbetter project to help raise awareness to wasteful model use and provide the community with an alternative.

please let me know if you have issues testing the model on your machine. i'd be very grateful to hear any feedback you might have!

Replies chronicallyoffl · 1d
Open on HN
Loading the discussion…

Domain filters

Stories from these domains are hidden from every list. Subdomains match too: blocking substack.com also hides danluu.substack.com.

    Help

    Keyboard

    j / k
    Move down and up the story list. The arrow keys scroll whatever has focus.
    Enter
    Read the marked story in the article reader.
    ]
    Read the next story in the same article-reader applet. Back returns to the previous story.
    p
    Pin or unpin the marked story, which keeps it in Pinned.
    n / N
    Move to the next or previous top-level comment in the window in front.
    c
    Collapse or expand that comment.
    f
    Hide or show the story list.
    Esc
    Close a menu or this help.
    Access key m
    Go to the menu bar. Most browsers take it with Alt on Windows and Linux, and Safari with Control and Option.
    ?
    Show this help.

    Windows

    Each story opens in a window holding its article above its discussion; drag the bar between them to share the room differently. A window can be moved by its title bar, resized from any edge, snapped to a half or a corner by dragging it there, maximised, or minimised to the bar at the foot of the page. Use Window > New reader window to open an empty reader, or a story row's Open in new reader window button to compare articles. Docked readers keep their articles when you select another story from the sidebar. Minimized readers can be restored and reused for their site. A window's Next story link reads on down the list in the same window.

    A link in a comment or an article to another Hacker News or Lobsters thread opens that thread in a window too. A link to a single HN comment opens the comment above its replies.

    While a story's window is in front, the Story and Discussion menus in the menu bar hold its commands: pinning, Next story, sorting, collapsing every thread, jumping to the first new comment. Each window also remembers where you were in its article and discussion, so a reload, or Back to a story that Next took you past, finds your place again. Closing a window forgets it.

    The whole arrangement lives in the address, so a bookmark or a shared link brings it back, and Back undoes the last change. Moving between Hacker News, Lobsters, their lists, Pinned and Find changes only the list, and leaves the windows open.

    The list

    The pin at the start of a row keeps the story in Pinned, and the cross at its end hides it. Pinned can be narrowed by words in the title, site or author, by source, and to the stories you haven't opened yet, and ordered by when you pinned them, by points or by comments; the filters are part of the address, so a filtered view can be bookmarked. Scroll past the end of the list to load more. Domain filters, in the View menu, hide every story from a site.

    Browsing view

    View > Windowed and View > Classic select the browsing view and save your default in this browser. Window view reuses a reader for each feed. Classic view opens stories and applets as pages. Open as a page and Open in a window are one-off actions that do not change your saved default. Direct page links always open as pages.

    Applets

    The Applets menu in the menu bar holds three tools, each a window of its own. Replies to me takes your Hacker News user name and lists the replies to your last thirty comments and stories, checking again every three minutes while it is open, and marking what is new since you last marked them read. Look up a user opens a profile on Hacker News or Lobsters, with their submissions and recent comments, as a commenter's name in any discussion does; the bar at the top of a profile looks up someone else in the same window, and Back returns to the one before. Who is hiring? filters the posts of HN's monthly hiring threads by the words you type.

    They read only what the sites publish to everyone, so none of them asks for a login, and your user name stays in this browser.

    Find

    Find takes any link and lists every time it was submitted to Hacker News and Lobsters, so you can read each discussion of it.

    About

    YAVCHN never sees your Hacker News or Lobsters login. The discussion is fetched from each site's public API; to vote or reply, follow the link above the discussion, or the arrow beside a comment, to the source's own site. Pins, hidden stories, filters and layout are kept in this browser only.

    Open source: github.com/paulmooreparks/yavchn. Built with PUDL.