pebble · machine intelligence lab · san francisco

Machines learn by doing.

We build your reinforcement learning environment. Reinforcement learning is an old idea. You improve at a trade by practicing it, and by knowing how each attempt turned out. Models improve the same way.

We build with you, around the capabilities you need and the outcomes you want, each one verifiable. Your machines improve on your work specifically, and the improvement is yours to keep.

Industry

What an environment is.

An environment wraps software around your work. A machine does the steps on your systems, and a check you approved grades each one.

Machines improve at whatever gets graded, so a machine in your environment improves at your process. The graded runs become training data, and a model trained on them has practiced your work toward the outcomes you want.

Engagement

Why you'd want one.

We build one of these around your process. Engagements run in three stages, and most customers start at the first and decide later how far to go.

01run

Working software.

We scope one process and build its environment, then hand it over doing the job, on a frontier model to start. The step map and the first checks get built here; the later stages reuse both.

02prove

A full audit trail.

The rest of the checks come online and every run is logged against them. Postmortems start from a step number, and reliability is a percentage you read off the run history.

03own

A model and a dataset you own.

An environment with checks is a training ground. We run an open model through it at volume and train on the checked outcomes, reinforcement learning on your own process. The model that comes out is tuned to the work and costs less per run than a frontier API. The run records are plain data, and when a stronger base model ships, they train its replacement.

Process

What working with us looks like.

1scope

You describe the process and the outcome you need from it. If an environment is the wrong tool for the job, we say so up front.

2discover

About a week inside your systems, under your access model, finding where checks can attach. We keep the step map and none of your data, and you get a written proposal with the end goal, the work packages, and the price.

3build

The environment and its checks, exercised against real runs of your process before we call it done.

4run

Handoff, on your infrastructure or ours, with a retainer for maintenance and the next stage when you want it.

Research

We also run a research program.

Pebble publishes its own research on matrix-valued state and measurable reasoning in language models. Findings go out as falsifiable notes with code and eval data. Instrument bugs are disclosed, and negative results stay published.

Read the research log →

Academic

Try Rockie.

Rockie is an RL environment we maintain for the public good. It runs experiment loops for academic groups doing HPC research, including ML. The system behind it is written up in our white paper, Scaling Experimentation, on making the full experiment loop run as software.

Rockie is tuned for scientific discovery, and we keep it public because an environment improves with every kind of work that runs through it. The hope is that the models people run in there leave human knowledge further along than they found it. An account takes a minute at rockielab.com.

Contact.

intro@pebbleml.com