We build your reinforcement learning environment. Reinforcement learning is an old idea. You improve at a trade by practicing it, and by knowing how each attempt turned out. Models improve the same way.
We build with you, around the capabilities you need and the outcomes you want, each one verifiable. Your machines improve on your work specifically, and the improvement is yours to keep.
An environment wraps software around your work. A machine does the steps on your systems, and a check you approved grades each one.
Machines improve at whatever gets graded, so a machine in your environment improves at your process. The graded runs become training data, and a model trained on them has practiced your work toward the outcomes you want.
every run leaves an artifact, and the model reads the whole shelf at once.
We scope the outcome you need and build the environment that reaches it, however many processes that touches. It ships doing the job, on a frontier model to start.
Every run is logged against the checks. Postmortems start from a step number, and reliability is a percentage read off the run history.
We train an open model on the runs your checks approved. It comes out tuned to your work and cheaper per run than a frontier API. The records can retrain whatever base model comes next.
You describe the process and the outcome you need from it. If an environment is the wrong tool for the job, we say so up front.
About a week inside your systems, under your access model, finding where checks can attach. We keep the step map and none of your data, and you get a written proposal with the end goal, the work packages, and the price.
The environment and its checks, exercised against real runs of your process before we call it done.
Handoff, on your infrastructure or ours, with a retainer for maintenance and the next stage when you want it.
Pebble publishes its own research on matrix-valued state and measurable reasoning in language models. Findings go out as falsifiable notes with code and eval data. Instrument bugs are disclosed, and negative results stay published.
Rockie is an RL environment we maintain for the public good. It runs experiment loops for academic groups doing HPC research, including ML. The system behind it is written up in our white paper, Scaling Experimentation, on making the full experiment loop run as software.
Rockie is tuned for scientific discovery, and we keep it public because an environment improves with every kind of work that runs through it. The hope is that the models people run in there leave human knowledge further along than they found it. An account takes a minute at rockielab.com.