> ## Documentation Index
> Fetch the complete documentation index at: https://langchain-5e9cc07a-preview-docsfi-1782565220-3e5e53a.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Run evaluations with Harbor

> Run Harbor evaluations and rollouts on LangSmith sandboxes with the harbor[langsmith] extra.

[Harbor](https://harborframework.com/docs) is a framework for evaluating and optimizing agents and language models in sandboxed environments, from the creators of [Terminal-Bench](https://www.tbench.ai). Harbor runs each trial in an isolated container, so you can parallelize evaluations and rollouts across many environments at once.

The `langsmith` Harbor environment runs those trials on LangSmith sandboxes. Select it with `-e langsmith` to execute Harbor jobs on LangSmith infrastructure, alongside providers such as AgentCore, Daytona, E2B, and Modal.

## Prerequisites

* A [LangSmith account](https://smith.langchain.com?utm_source=docs\&utm_medium=cta\&utm_campaign=langsmith-signup\&utm_content=langsmith-sandbox-harbor) and an API key.
* Python with `pip`.

## Install

Install Harbor with the `langsmith` extra:

```bash theme={null}
pip install "harbor[langsmith]"
```

## Authenticate

Harbor authenticates with your LangSmith credentials. Set an API key:

```bash theme={null}
export LANGSMITH_API_KEY="<LANGSMITH_API_KEY>"
```

`LANGCHAIN_API_KEY` works as well. Alternatively, select a [LangSmith SDK profile](/langsmith/profile-configuration) instead of exporting a key:

```bash theme={null}
export LANGSMITH_PROFILE=prod
```

## Run an evaluation

Run a Harbor job and select the LangSmith environment with `-e langsmith`:

```bash theme={null}
harbor run -d "<org/name>" \
  -m "<model>" \
  -a "<agent>" \
  -e langsmith \
  -n "<n-parallel-trials>"
```

Harbor creates one LangSmith sandbox per trial, runs the agent and verifier inside it, then tears the sandbox down when the trial finishes.

## Configure the sandbox environment

The LangSmith environment boots each sandbox from a filesystem snapshot. Provide one of the following in your Harbor task:

* **Prebuilt image**: set `[environment].docker_image` in `task.toml`. Harbor reuses or creates a snapshot from that image.
* **Existing snapshot**: pass `environment.kwargs.snapshot_name` to boot from a [snapshot](/langsmith/sandbox-snapshots) you already created.
* **Dockerfile**: include an `environment/Dockerfile`. Harbor builds a snapshot from it with the [build-from-Dockerfile flow](/langsmith/sandbox-snapshots#build-a-snapshot-from-a-dockerfile), using the task `environment/` directory as the build context.

Tune the sandbox lifecycle with environment kwargs, passed on the command line with `--ek`:

```bash theme={null}
harbor run -d "<org/name>" \
  -m "<model>" \
  -a "<agent>" \
  -e langsmith \
  -n "<n-parallel-trials>" \
  --ek idle_ttl_seconds=0 \
  --ek delete_after_stop_seconds=7200
```

* `idle_ttl_seconds`: stops an idle sandbox after this many seconds. Set `0` to disable the idle timeout.
* `delete_after_stop_seconds`: deletes a stopped sandbox after this many seconds.

## Run Deep Agents on LangSmith

[Deep Agents](/oss/python/deepagents/overview) run on LangSmith sandboxes through Harbor's built-in `langgraph` agent and the `-e langsmith` environment. The full setup: staging the Deep Agents packages, the LangGraph project (`langgraph.json` with the `deepagent` graph), and the exact `harbor run` command, is maintained in the Deep Agents eval guide:

→ [Running Deep Agents on Harbor / Terminal Bench 2.0](https://github.com/langchain-ai/deepagents/blob/main/libs/evals/CONTRIBUTING.md#harbor--terminal-bench-20)

The [LangSmith-specific options](#configure-the-sandbox-environment) (authentication, `-e langsmith`, and the `--ek` sandbox-lifecycle kwargs) apply to those runs as well.

## Multi-container tasks

The LangSmith environment supports multi-container tasks. Include an `environment/docker-compose.yaml` file in your task definition to run several containers per trial. See the [Harbor sandbox documentation](https://harborframework.com/docs/run-jobs/cloud-sandboxes) for details.

***

<div className="source-links">
  <Callout icon="terminal-2">
    [Connect these docs](/use-these-docs) to Claude, VSCode, and more via MCP for real-time answers.
  </Callout>

  <Callout icon="edit">
    [Edit this page on GitHub](https://github.com/langchain-ai/docs/edit/main/src/langsmith/sandbox-harbor.mdx) or [file an issue](https://github.com/langchain-ai/docs/issues/new/choose).
  </Callout>
</div>
