Hacker Newsnew | past | comments | ask | show | jobs | submit | megadragon9's commentslogin

this makes me wonder, does having kids change the list of things on your list? and why?


No, it just means you need a lot more money because you want the things on your list for your kids also.


Having children changes pretty much everything in your life.


I think it's more about the scar tissue (a.k.a. intuition) when building LLMs from scratch than whether it's transferable to job search. Maybe the person will decide they don't like LLMs and not develop that into a career, or maybe they become a researcher in another field because that's a better way to solve problems inherent in LLMs.


I did something similar but with smart home devices. I use the homebridge interface to connect my smart home devices to Apple's homekit protocol. Some homebridge plugins for my devices were outdated and no longer maintained, so I asked Codex/Claude to help me create a patch of it as a local fork, so my smart home devices can still run without problems.

It does feel magical when these agents can debug in the real-world, like turning on/off my living room lights and using another living room camera to take a snapshot of the living room to see whether it worked or not.


i created something similar a while back. Inspired by micrograd for showing the connection between math/calculus and code, then building along the way to a full NumPy deep learning library that I pretrained GPT-2 124M model with it. One way to learn is to trace through the PRs merged to the repo in chronological order.

The philosophy is the same: "What I cannot create, I do not understand." https://github.com/workofart/ml-by-hand


wow this is nice work. i am diving into this world and was hoping to this exact thing. i am going to look to see how you did it. ty for sharing.


I did something similar. It started off as a "self-improving agent" project, inspired by autoresearch, then later on I reframed it as "harness training" (discrete program search) borrowing the mental model from ML training.

I "trained" the harness on a subset of Terminal-Bench 2.0 tasks while keeping the LLM (local Qwen3.6-35B A3B) frozen. Making LLM inference and the task environments fully deterministic was necessary for clean credit assignment. I learned this the hard way after spending the initial 1 month on experiment noise.

My final results showed that on the full 89-task Terminal-Bench 2.0 suite, the trained harness matched or beat the official Terminus 2 harness for four LLMs that it never collaborated with during training (e.g. GPT-OSS-120B score increased from 18.7% to 36%, while using 55% fewer input tokens per solve). A harness trained only on SWE-bench improved Terminal-Bench scores too. Here's the write-up: https://www.henrypan.com/blog/2026-07-18-harness-training/

I packaged the training loop as a PyTorch-style framework. https://github.com/workofart/harness-training


I worked on this project (https://github.com/workofart/harness-training) for the past few months to reframe "Agent-driven Self-improving Harness" to "Harness Training".

The idea is simple, the harness is trained once with a frozen task LLM against a given task environment. Then you can then swap out the task LLM to any model and evaluate the "frozen trained harness" with any task LLM on any new task environment.

Since this was a general problem, I took the chance to create a general PyTorch-like training framework. Right now, you can train with any OpenAI-compatible API for interfacing with the task LLM and train against Terminal-Bench or SWE-Bench tasks, but you can easily extend it to support any task environments.

I wrote a blog post (https://www.henrypan.com/blog/2026-07-18-harness-training) on this journey, including (but not limited to):

- results from using this harness training framework to improve general capabilities across many task LLMs to beat Terminal Bench 2.0 (Terminus Harness) and also transfer learnings towards better task-solving abilities in unseen task environments (e.g. harness trained on SWE-Bench tasks solving Terminal Bench tasks)

- how this framework is built

- learnings on what was missing in my initial version of the project (hint: determinism)


Interesting project. Do you think manual memory management help understand computational graph lifecycle better, or does it distract from backprop itself?

btw, I went down the micrograd path with numpy-primitives all the way to building a PyTorch clone that can pre-train and post-train LLMs (https://github.com/workofart/ml-by-hand). My learning focus was on the math/calculus <-> high-level APIs, instead of efficiency. I'm glad to see more people tackling this problem from different angles.


ngl, it distracts from backprop itself a little, but teaches a lot about memory management. I did it this way because in parallel I wanted to get better at C, but if your aim is to purely work on ML fundamentals, it’s probably better to do it in python


Model built and trained using a hand-built deep learning library (numpy primitives)


I'm continuing to expand my own deep learning library [1] built with numpy-primitives to support LLM post-training techniques like supervised fine-tuning (SFT) and reinforcement learning with GRPO. It's a good learning experience to work without all the high-level abstractions to "build a wheel" and "use that wheel to build a car".

I'm also looking into coding harness self-improvement [2]. An inner LLM (raw LLM request) + harness solves coding tasks, an outer agent like Claude or Codex that proposes harness changes. I experimented with many things in the past few months that made me realize this self-improvement thing that everyone is talking about is just an experiment design problem. I wrote about it here [3]. I'm continuing to improve the infra around the self-improvement loop, to increase signal-to-noise ratio per experiment. I'm also generalizing the infra to expand beyond terminal bench tasks and to collect some data across different models (harness-bound vs model-bound).

[1] https://github.com/workofart/ml-by-hand

[2] https://github.com/workofart/harness-experiment

[3] https://www.henrypan.com/blog/2026-05-25-self-improvement-ha...


looks like elon web services (EWS) is the master plan all along :D


Everything will run on Elon's Costly Compute servers.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: