We build your reinforcement learning environment. Reinforcement learning is an old idea. You improve at a trade by practicing it, and by knowing how each attempt turned out. Models improve the same way.
We build with you, around the capabilities you need and the outcomes you want, each one verifiable. Your machines improve on your work specifically, and the improvement is yours to keep.
An environment wraps software around your work. A machine does the steps on your systems, and a check you approved grades each one.
Machines improve at whatever gets graded, so a machine in your environment improves at your process. The graded runs become training data, and a model trained on them has practiced your work toward the outcomes you want.
We build one of these around your process. Engagements run in three stages, and most customers start at the first and decide later how far to go.
We scope one process and build its environment, then hand it over doing the job, on a frontier model to start. The step map and the first checks get built here; the later stages reuse both.
The rest of the checks come online and every run is logged against them. Postmortems start from a step number, and reliability is a percentage you read off the run history.
An environment with checks is a training ground. We run an open model through it at volume and train on the checked outcomes, reinforcement learning on your own process. The model that comes out is tuned to the work and costs less per run than a frontier API. The run records are plain data, and when a stronger base model ships, they train its replacement.
You describe the process and the outcome you need from it. If an environment is the wrong tool for the job, we say so up front.
About a week inside your systems, under your access model, finding where checks can attach. We keep the step map and none of your data, and you get a written proposal with the end goal, the work packages, and the price.
The environment and its checks, exercised against real runs of your process before we call it done.
Handoff, on your infrastructure or ours, with a retainer for maintenance and the next stage when you want it.
Pebble publishes its own research on matrix-valued state and measurable reasoning in language models. Findings go out as falsifiable notes with code and eval data. Instrument bugs are disclosed, and negative results stay published.
Rockie is an RL environment we maintain for the public good. It runs experiment loops for academic groups doing HPC research, including ML. The system behind it is written up in our white paper, Scaling Experimentation, on making the full experiment loop run as software.
Rockie is tuned for scientific discovery, and we keep it public because an environment improves with every kind of work that runs through it. The hope is that the models people run in there leave human knowledge further along than they found it. An account takes a minute at rockielab.com.