Hi, this is the research project I'm working on. I must say, building a GUI harness is way more complicated than I thought. The goal is something like Minecraft. You know, a few simple tools that a user or an LLM can use to build large & complex worlds(aka use-cases).
Also, I want to build something that will have all my software in one place(independent from cloud) and goes beyond chat interface.
Few notes:
#Easy to run:
- Just download the single exe file and double click on it. No installation, docker, etc.
- Written in go-lang.
- ~40MB file size for CPU(no CUDA compiled in) version.
- 70-100MB RAM usage. Before LLM starts kicking of course.
It doesn’t use a web browser. It has it’s own text and layout engines build from scratch - basically fancy SVG. Everything is compiled into a one bin file, including:
- code interpreter
- llama.cpp
- whisper.cpp
- omnivoice.cpp
- and more
#GUI:
- User talks to it and the LLM generates code, which is executed and outputs vector components(kind of like .SVG).
- The code can use only 7 simple vector graphics functions(Rectangle, Circle, Line, Image, Text, List, Div). Each has a few parameters.
- Everything is made from them - Button, Slider, Color picker, Calendar, etc. It's fully hackable. Users can even create new never-seen GUI components.
In other words: Every(!) pixel on the screen is vector graphics, which was output from the code generated by LLM.
#Storage:
- Build from scratch(no OS files & folders).
- It records every change! Revert your mistakes. Revert LLM mistakes.
- It started as a boring feature, but I quickly realized how often I use the back button. It takes UX to another level.
- You don't have to keep the full history, just keep the last hour, day, or week.
In other words: If you clicked on something and regret it or the agent went wild, you can easily bypass the action like it never happened.
#LLMs:
- Testing gemma-31b@fp8, Qwen-27b@fp8 and DeepSeek-V4-flash@fp8. I hope soon, it will be possible to go sub 20B.
- A big advantage of open-weight models is they show reasoning, which helped me in so many cases expose “bugs” in context.
- Context enginnering is probably 50% there. So many improvements needed.
- Optionally, It can use models from OpenRouter to test SOTA models.
#Brush:
- Layouts are very minimalistic. For example, there are no Delete or Rename buttons for navigation on left side of the screen.
- Brush is non-verbal part of the communication.
- Hold CTRL key and brush with your mouse/touchpad.
- You can selected Div(part of layout) or select grid on a Div.
- Agent doesn't see a screen, it edits code and knowledge base.
In other words: Minimalism is great for UX experience and executing second level features is done by pointing on a part(s) of the screen and talking.
Check the video.
This is just beginning. Tons of stuff need to be improved, including my presentation skills. Overall, It's stable(no AI slop). Waiting for agent to finish is currently biggest UX problem which I'm solving by improving context enginnering.
Feel free to ask any questions.
submitted by
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.