Environments
An Agor environment is a set of commands attached to a branch. You supply the commands to start, stop, reset, and inspect your app; Agor runs them when you use the branch’s environment controls. You choose where the app runs.
Docker Compose on a development machine or shared server is the most
straightforward starting point: the commands and predictable URLs fit entirely
in .agor.yml. A remote provider uses the same interface, sometimes with a small
script that calls its API, waits for startup, or discovers the app URL. Agor does
not require Docker, provision infrastructure on your behalf, or ship a native
integration for every provider: if you can write a bounded command that triggers
your environment, you can connect it.
A repo ships a library of named variants (lean, postgres, full, …) that
every branch can pick from. Variants are Handlebars-templated and get rendered
into a per-branch snapshot that the Start/Stop/Nuke/Logs buttons execute
against.
What it looks like on the card: a green env pill, one-click open, and start/stop/logs/edit icons right next to your branch’s PR badge.
Quick start: Docker Compose
This example assumes a self-hosted Agor instance whose environment executor can access Docker, Compose, and the branch checkout. On a shared box, the operator must deliberately grant that access: control of a Docker engine is powerful, not a substitute for filesystem or user isolation.
Suppose your repo already has a compose.yml that builds your app, serves HTTP
on container port 3000, and exposes /health. Its port mapping should accept a
host port rather than hard-code one:
# compose.yml — adapt your existing app service
services:
app:
build: .
ports:
- '${APP_PORT}:3000'Put this complete environment definition in the repo root:
# .agor.yml
version: 2
environment:
default: local
variants:
local:
description: Docker Compose on this Agor host
start: >-
APP_PORT={{add 10000 branch.unique_id}}
docker compose -f compose.yml -p agor-{{branch.id}} up -d --build
stop: >-
APP_PORT={{add 10000 branch.unique_id}}
docker compose -f compose.yml -p agor-{{branch.id}} down
nuke: >-
APP_PORT={{add 10000 branch.unique_id}}
docker compose -f compose.yml -p agor-{{branch.id}} down -v
logs: >-
APP_PORT={{add 10000 branch.unique_id}}
docker compose -f compose.yml -p agor-{{branch.id}} logs --tail=100
health: http://localhost:{{add 10000 branch.unique_id}}/health
app: http://{{host.ip_address}}:{{add 10000 branch.unique_id}}Each branch gets its own Compose project and deterministic port, so there is no
URL-discovery script or provider binding to maintain. Choose a free port range
on your host and keep the resulting ports below 65536. Avoid fixed
container_name values or shared explicit volume names in your Compose file,
which would defeat per-project separation.
health is the URL the Agor daemon probes; app is the URL your
browser opens. This example assumes the daemon can reach the Docker host at
localhost; if not, use a daemon-reachable address for health. Set
host.ip_address to an address your browser can reach
(not necessarily a public IP). Docker publishes the app on the host’s
interfaces, so restrict access with your network policy and app authentication.
To connect the controls:
- A repo administrator imports
.agor.ymlin the repository environment editor. - In the branch’s Environment tab, select local and Render the commands. Review the resolved project name, ports, and URLs.
- Select Start, inspect the output and health, then open the app link.
- Use Logs for a bounded log tail and Stop when finished.
- Use Nuke only when you also intend to remove the project’s Compose-managed volumes. External volumes and bind-mounted files are not erased by this command.
| Field | In this example |
|---|---|
start | Build and start the branch’s Compose project, then exit. |
stop | Remove its containers/network, retaining named volumes. |
nuke | Also remove its Compose-managed volumes; destructive and optional. |
logs | Return recent output and exit; optional. |
health | An HTTP URL, not a health command; optional. |
app | A browser link; optional. |
A successful Start means the command succeeded, not that the app is ready. With a health URL, Agor observes HTTP health separately; without one, a successful Start has unknown health. This example rebuilds from the local checkout, including uncommitted changes; live reload additionally requires your own Compose mounts and app watch configuration.
Where your app runs
Self-hosted machines may run Compose locally. Agor Cloud does not provide a local Docker host for your app: use commands that drive a remote provider where your deployment enables them, or its configured webhook surface. Executor tools, credentials, network access, and time limits still apply, and an app launched inside a Cloud executor does not receive public ingress.
Overview
Configuration layers
Managed environments are built out of three layers, precedence low → high:
┌────────────────────────────────────────────────────────┐
│ 1. .agor.yml — repo-shared, committed to git │
│ environment.variants.{lean,postgres,full,…} │
└───────────────────────────┬────────────────────────────┘
│ imported into repo.environment
▼
┌────────────────────────────────────────────────────────┐
│ 2. repo.template_overrides — DB-only, per deployment │
│ host.ip_address, custom.internal_registry, … │
└───────────────────────────┬────────────────────────────┘
│ deep-merged into render context
▼
┌────────────────────────────────────────────────────────┐
│ 3. branch rendered snapshot — plain strings │
│ start_command, stop_command, nuke_command, │
│ logs_command, health_check_url, app_url │
└────────────────────────────────────────────────────────┘- Layer 1 is authored by repo maintainers and lives in
.agor.yml. It’s portable: clone the repo,Import .agor.yml, done. - Layer 2 is where each Agor deployment pins its own values (host IP, internal registry, AWS profile) without touching the shared file. Stored in the DB, never exported.
- Layer 3 is the output: a flat set of rendered strings attached to the branch. Executors run these verbatim, with no templating at execute time.
Re-rendering is explicit. Editing a variant or an override does not retroactively rewrite existing branches. You hit the Render button when you want a branch to pick up changes.
Branch Environment tab, with variant YAML on top, picker in the middle, rendered snapshot below. The extends: sqlite line on the postgres variant shows variants can inherit from each other.
Picking a variant. This historical screenshot predates the rename from full to rich. full remains a deprecated compatibility alias; HA is still a separate variant.
Template variables
Every field in a variant is a Handlebars template. The render context is built in packages/core/src/templates/handlebars-helpers.ts.
Variables
| Variable | Source | Example | Notes |
|---|---|---|---|
{{branch.unique_id}} | Branch | 3 | Auto-incrementing integer. Use for deterministic port offsets. |
{{branch.id}} | Branch | 01a0da75-7bb6-… | Stable UUID for the Agor branch. Good for per-branch resource names and provider bindings. |
{{branch.name}} | Branch | feat-new-filter | Slugified branch name. Good for docker compose -p. |
{{branch.ref}} | Branch | feat-new-filter | Exact Git ref recorded for the branch. Use it when a remote provider builds from a pushed ref. |
{{branch.path}} | Branch | /home/you/agor/feat-new-filter | Absolute path on disk. |
{{branch.base_ref}} | Branch | main | Source branch/tag name the branch was created from (the “Base Branch”/“Base Tag” in the create dialog). The name, not a SHA. Empty string if unknown. |
{{branch.ref_type}} | Branch | branch | branch or tag, indicating whether base_ref names a branch or a tag. Defaults to branch. |
{{repo.slug}} | Repo | superset | Agor’s local repository slug, used in managed paths and UI. It is not guaranteed to include a provider owner. |
{{repo.github_slug}} | Repo | apache/superset | Credential-free GitHub owner/repository identity derived from the registered github.com remote. Empty for non-GitHub or unknown remotes. |
{{host.ip_address}} | Auto-detected | 10.0.1.42 | Fallback order: template_overrides.host.ip_address → daemon.host_ip_address in ~/.agor/config.yaml → autodetect. |
{{custom.<key>}} | Branch custom context | {{custom.compose_profile}} | Anything you store in branch.custom_context. |
{{custom.<key>}} via overrides | template_overrides.custom.* | {{custom.internal_registry}} | Repo-wide fallback if the branch doesn’t set a key. |
Math helpers
| Helper | Example | Result (with unique_id = 3) |
|---|---|---|
add | {{add 9000 branch.unique_id}} | 9003 |
sub | {{sub 9000 branch.unique_id}} | 8997 |
mul | {{mul branch.unique_id 10}} | 30 |
div | {{div 60 branch.unique_id}} | 20 |
mod | {{mod branch.unique_id 4}} | 3 |
Math helpers are deterministic. The same branch always renders the same ports, which is what keeps parallel branches from colliding.
Use {{shellQuote value}} for any rendered string embedded in a shell command: it quotes the value as one POSIX shell argument instead of treating branch text as shell syntax.
Broken template syntax is reported at render time, not at save time. Click Render to see the exact error with a line number.
The .agor.yml file
.agor.yml lives at the branch root ($BRANCH_PATH/.agor.yml) and is the shared, committed description of a repo’s environment variants. It’s a regular repo file: you edit it on a branch, open a PR, and it rolls out when merged.
Schema
version: 2
environment:
default: lean # which variant new branches get
variants:
<variant-name>:
description: 'Human-readable one-liner'
extends: <base-variant> # optional, single-level only
start: '<shell command template>'
stop: '<shell command template>'
nuke: '<shell command template>' # optional
logs: '<shell command template>' # optional
health: '<url template>' # optional
app: '<url template>' # optionalstart, stop, nuke, and logs are rendered lifecycle fields. In the
default hybrid mode they can be shell commands or explicit http:// /
https:// webhooks. Webhooks are invoked with HTTP GET. health and app are
always URL templates, not shell commands.
Required fields
version: 2at the top level.environment.defaultmust reference a variant that exists inenvironment.variants.- Each variant (after
extendsresolution) needs at leaststartandstop.
nuke, logs, health, and app are optional. The corresponding UI buttons simply hide or no-op when absent.
Runtime URLs from Start
Some providers choose a different URL each time they create or resume an
environment. A successful start may therefore return this deliberately tiny,
optional JSON object:
{ "app": "https://preview.example.test", "health": "https://preview.example.test/health" }Both keys are optional, so {} is valid. Unknown keys, non-HTTP(S) URLs,
embedded credentials, query strings, and fragments are rejected. Omitted keys
fall back to the rendered static app or health value. Runtime values last
until the next Start/Stop/Nuke lifecycle boundary; Start is the reconciliation
point rather than a permanent provider-name lookup.
- A shell Start emits exactly one line:
AGOR_ENVIRONMENT_RESULT={"app":"https://...","health":"https://..."}. Agor removes that control line from captured command output. - A webhook Start returns the same object with an
application/json(orapplication/*+json) content type. Non-JSON webhook bodies retain their legacy meaning and are ignored as lifecycle results.
app is an optimistic browser link and may point to an authenticated/private
provider page. health is different: the daemon polls it continuously. A
runtime health URL must resolve only to public IP space, is connected through a
DNS-pinned no-redirect request, and should be returned only when the endpoint is
actually reachable by the Agor daemon. Static operator-authored health URLs may
still target localhost or private infrastructure.
When Start completes without either a runtime or static health URL, the environment becomes Started; health unavailable rather than remaining on a spinner or claiming to be healthy. The App link remains available when known.
extends (single-level only)
A variant may extend one base variant. The base variant itself must not have an extends key. Chains (full → postgres → lean) are rejected at save time with:
variant 'full' extends 'postgres', which itself extends 'lean' —
chains deeper than one level are not allowedResolution is a straight per-field merge: child fields win, omitted fields inherit from the base.
Worked example for Superset: lean / postgres / full
version: 2
environment:
default: lean
variants:
lean:
description: 'SQLite-backed, single-container, fast iteration'
start: 'docker compose -f docker-compose-light.yml -p agor-{{branch.name}} up -d'
stop: 'docker compose -f docker-compose-light.yml -p agor-{{branch.name}} down'
nuke: 'docker compose -f docker-compose-light.yml -p agor-{{branch.name}} down -v'
logs: 'docker compose -f docker-compose-light.yml -p agor-{{branch.name}} logs --tail=100'
health: 'http://{{host.ip_address}}:{{add 9000 branch.unique_id}}/health'
app: 'http://{{host.ip_address}}:{{add 5000 branch.unique_id}}'
postgres:
description: 'Postgres + Redis + Celery — closer to prod'
extends: lean # inherits health + app from lean
start: 'docker compose -p agor-{{branch.name}} up -d --build'
stop: 'docker compose -p agor-{{branch.name}} down'
nuke: 'docker compose -p agor-{{branch.name}} down -v'
logs: 'docker compose -p agor-{{branch.name}} logs --tail=100'
full:
description: 'Postgres + Redis + Celery + worker + beat'
extends: lean # NOT extends: postgres — chains are not allowed
start: 'COMPOSE_PROFILES=full docker compose -p agor-{{branch.name}} up -d --build'
stop: 'docker compose -p agor-{{branch.name}} down'
nuke: 'docker compose -p agor-{{branch.name}} down -v'
logs: 'docker compose -p agor-{{branch.name}} logs --tail=100'Notice that postgres and full both extend lean (not each other). That’s the single-level rule. Both inherit health and app from lean and override the rest.
Agent configuration workflow
Agents use the same environment service as the branch card’s Play/Stop controls.
Discover the tools with agor_search_tools and inspect their schemas with
agor_get_tool_details before calling them:
- Inspect the target repository’s
.agor.yml, the imported definitions fromagor_repos_get, and the branch’s current rendered commands and lifecycle state. For a remote provider, adapt the provider-neutral lifecycle pattern and Railway reference to your application. - Review the complete configuration, then call
agor_repos_import_environmentwithrepoIdand the sourcebranchId. This requires admin access and branch filesystem read access. It replaces the repository’s variants and default, preserves deployment-local overrides, and leaves existing branch snapshots alone. Editing the file by itself does not apply it. The admin Import.agor.ymlaction below is the fallback when the tool is unavailable. - Call
agor_environment_setwith the targetbranchIdand desiredvariant. Omitvariantto re-render the current one. Stop an active environment before switching variants, and inspect the resulting commands before starting. - Call
agor_environment_start, then inspectagor_environment_healthandagor_environment_logs. An accepted request does not prove readiness. - Use the reported app URL, and call
agor_environment_stopwhen finished. Useagor_environment_nukeonly when destructive cleanup is intended.
Import changes repository-wide definitions, including those other branches may
render later. Keep every intended variant in the replacement and inspect existing
configuration before changing it. Environment commands use the invoking user’s
Global environment variables; collect missing credentials through
agor_widgets_request_env_vars, not YAML or rendered command arguments.
Import / export
Two admin-only actions in the repo editor header:
| Action | Effect |
|---|---|
Import .agor.yml | Parses $BRANCH_PATH/.agor.yml and replaces environment.variants + environment.default in the DB. template_overrides and branch snapshots are untouched. |
Export .agor.yml | Writes the current environment.variants + environment.default to $BRANCH_PATH/.agor.yml. template_overrides is stripped before writing. |
Both paths go through a confirm dialog. Import is replace, not merge. A variant that exists in the DB but not in the file gets dropped.
Because the target is always the file in the currently open branch, the normal workflow is:
- Iterate on
.agor.ymlon a branch. - Export (or hand-edit) the file.
- Commit and PR it like any other repo change.
- On merge, other branches can
Importon their next sync.
Deployment-local overrides
template_overrides is a per-repo, per-deployment block that lives in the DB, not in .agor.yml. It’s where each Agor instance pins the concrete values that the shared variants reference (IPs, registries, profile names).
Why it exists
Superset’s upstream .agor.yml shouldn’t hard-code 10.0.1.42 as the host IP. That’s Preset’s infra, not Apache’s. Instead, Superset’s variants reference {{host.ip_address}} and Preset sets it once, per-repo, in template_overrides. The shared file stays clean; the deployment-specific values stay in the deployment.
Schema
template_overrides:
host:
ip_address: '10.0.1.42' # overrides daemon config / autodetect
custom:
internal_registry: 'registry.preset.io'
aws_profile: 'preset-dev'
compose_profile: 'preset'The structure is a plain { host, custom } map. template_overrides is root-level and applies to all variants. There are no per-variant overrides.
Precedence
At render time the context is assembled in this order (later wins):
- Daemon defaults (autodetected
host.ip_address, etc.). repo.template_overrides(deep-merged).branch.custom_context(forcustom.*).- Branch identity (
branch.*,repo.*).
What it is not
⚠
template_overridesis not a secret store. It’s visible to every user with read access to the repo (admins edit, members read). Use it for infrastructure identifiers (IPs, registry hostnames, profile names), never for API keys, tokens, or passwords.For secrets, use your user-level environment variables (Settings → Environment Variables, Global scope). Environment commands receive the global variables of the user who presses Start/Stop/Nuke/Logs; session-selected variables are not passed to them.
The import/export parser refuses any template_overrides: key found in .agor.yml and the exporter always strips it. This is a defense-in-depth guarantee that values like host.ip_address: 10.0.1.42 cannot be accidentally committed upstream to a public repo.
Variants in the branch
Each branch has:
environment_variant, the chosen variant name (e.g."postgres").start_command,stop_command,nuke_command,logs_command,health_check_url,app_urlhold the rendered strings, fully substituted, no Handlebars left.
The Environment tab has two stacked editors:
| Editor | Who edits | What it shows |
|---|---|---|
| Repo editor (top) | Admins only | version, environment.variants, environment.default, template_overrides as YAML |
| Branch editor (bottom) | Admins only | The 6 rendered commands for this branch |
Members see both editors read-only. Users with branch all permission and admins can use the Variant picker and the Render button in between them.
Saving the repo YAML replaces its complete environment configuration: deleted variants, fields, and overrides are removed. Existing branch snapshots remain unchanged until Render. API updates that include environment use the same replacement semantics; omit that key to leave it unchanged.
The Render flow
[ Variant: postgres ▾ ] [ Render ▸ ]- Pick a variant from the dropdown.
- Click Render.
- Agor re-evaluates
variant + template_overridesand overwrites the branch snapshot.
If the snapshot has unsaved manual edits (admins only, since members can’t make them), Render prompts a confirm: “Rendering will discard your local edits. Continue?”
No auto-tracking
If an admin edits a variant after a branch has rendered, the branch snapshot does not update until someone clicks Render. There is no dirty indicator, no banner, no auto-rerender. Render explicitly when you are ready to apply the updated commands.
Empty state
A repo with no variants configured shows a disabled picker and:
No environment variants configured. Ask an admin to set up commands in the repo editor above.
Permissions
Two orthogonal axes govern access: who can edit what, and who can run Start/Stop/Nuke/Logs.
Edit permissions
| Action | Member | Admin |
|---|---|---|
| View repo config (top editor) | ✅ read-only | ✅ |
Edit environment.variants / environment.default | ❌ | ✅ |
Edit template_overrides | ❌ | ✅ |
Import / export .agor.yml | ❌ | ✅ |
| Pick variant + Render to branch | Requires branch all | ✅ |
| Hand-edit branch rendered commands | ❌ | ✅ |
The reason members cannot hand-edit rendered commands: admins curate the set of commands that can run (the variants), and users with branch all permission pick from that list. This makes the variant library the effective allowlist and complements the execution-time deny-list guard.
Execution permissions
Start, Stop, Restart, Nuke, Logs, and Render controls require branch all permission or admin access. Health/status reads can remain available to users who can view the branch because they do not run the configured shell/log commands.
Health-check lifecycle
Automatic health monitoring is active only while a non-archived environment is
starting or running. A healthy starting observation promotes the durable
state to running. Stop, failure, archive, and deletion remove the environment
from monitoring; daemon startup reconciliation likewise rediscovers only
non-archived starting and running rows.
The rendered health value is stored as branch configuration even when it is
not currently a usable URL. This keeps an inactive or optional managed
environment from blocking branch and session creation. Agor validates the URL
against its outbound-request policy immediately before an actual server-side
probe. An unsafe target is never contacted and is reported as unhealthy while
the environment is active.
An explicit health/status request does not turn monitoring back on for a
stopped, stopping, archived, or deleted environment. It may return a transient
diagnostic probe for an environment already in error, but that result does not
change durable lifecycle state. Active observations are fenced against
concurrent stop, archive, delete, restart, and health-URL changes, so a late
success cannot restore an older lifecycle state.
Instance guidance for operators
Operators can display a fixed informational notice at the top of every branch’s
Environment tab. Set this top-level key in ~/.agor/config.yaml and restart
the daemon (coordinate the same configuration across HA replicas):
environment_disclaimer_markdown: |
Commands can launch remote environments such as GitHub Codespaces or your own provider.
Agor runs your commands and reports their results; it does not host your application
in the executor. Servers started inside Cloud executor Pods are not exposed and
commands are time-limited. Your remote provider's security and billing apply.
See the [environment guide](https://agor.live/guide/environment-configuration)
for configuration and cleanup responsibilities.This is public, instance-wide policy content, not a secret store. Repositories, variants, and branch snapshots cannot override it. It supplements the dynamic capability summary; text cannot enable execution or change permissions.
Omit the key or use an empty string to show no notice. The limit is 4000 UTF-16 code units (most characters count as one; emoji may count as two). Supported Markdown is paragraphs, emphasis, lists, and absolute HTTP(S) documentation links without embedded credentials. Raw HTML, images, embeds, code blocks, and interactive plugins are not rendered. Unsafe and relative URLs are not clickable.
Migration from pre-variants Agor
No forced user action. The upgrade is silent:
| Before | After |
|---|---|
repos.environment_config: { up_command, down_command, … } | repos.environment: { version: 2, default: "default", variants: { default: <old value> }, template_overrides: {} } |
Branch start_command / stop_command / … | Unchanged. They’re already the rendered snapshot this design formalizes. |
.agor.yml v1 (flat environment: { start, stop, … }) | Still parses. Loaded as variants.default. Export writes v2 going forward. |
daemon.host_ip_address in ~/.agor/config.yaml | Still works as the fallback for {{host.ip_address}}. For new setups, prefer per-repo template_overrides.host.ip_address. |
Existing branches keep running against their existing rendered commands. Nothing re-renders until someone hits Render or creates a new branch.
Security notes
Two things worth internalizing before you ship a variant to the rest of the team.
1. Deny-list runs on the rendered string, at execute time
The deny-list guard (the thing that blocks destructive commands like rm -rf /) runs on the final rendered command, at execute time (not against the template, and not at save time). That means:
- A template can render to a denied command under some branch contexts and not others.
- The first visibility you have that a command is denied is when someone clicks Start / Stop / Nuke and it fails.
- Keep your templates simple. If the rendered shape of a command isn’t obvious from reading the template, add a
description:to the variant so teammates know what they’re launching.
2. template_overrides is not a secret store
Restating the callout from above because it’s the single most common mistake:
template_overridesvalues are visible to every user with repo access.- Values get baked into shell strings at render time. Anyone who can read the branch snapshot sees them.
- Use it for IPs, hostnames, registries, profile names. Not for API keys, tokens, or passwords.
- For secrets, use user-level Global environment variables. They are injected into the command process of the user who runs the action and never land in a rendered command. Session-selected variables are not available to environment commands.
3. Managed environment execution mode
Agor instances choose how lifecycle fields are handled. The default is
hybrid:
# ~/.agor/config.yaml
execution:
managed_envs_execution_mode: hybridIn hybrid mode, lifecycle fields can be shell commands or explicit
http:// / https:// webhooks. URL-shaped fields use webhook execution, so a
repo can use webhooks for one action and commands for another.
Some instances are configured for webhook-managed environments:
# ~/.agor/config.yaml
execution:
managed_envs_execution_mode: webhook-onlyIn webhook-only mode, rendered start, stop, nuke, and logs fields
must render to explicit http:// or https:// URLs. Agor sends a GET request
to the URL.
environment:
default: remote
variants:
remote:
start: 'https://orchestrator.example.com/agor/start?branch={{branch.name}}'
stop: 'https://orchestrator.example.com/agor/stop?branch={{branch.name}}'
logs: 'https://orchestrator.example.com/agor/logs?branch={{branch.name}}'
health: 'https://apps.example.com/{{branch.name}}/health'
app: 'https://apps.example.com/{{branch.name}}'V1 webhooks are intentionally simple: GET only, no custom headers, no request body, no signing, and redirects are not followed. V1 webhook targets must be public HTTP(S) destinations; localhost, private/link-local ranges, and URL credentials are rejected. Use a trusted orchestration service, and avoid secrets in query strings because rendered branch snapshots and infrastructure logs may expose them.
Bounded external commands in HA
HA hybrid mode supports remote-environment trigger scripts, not application
hosting. Agor dispatches a bounded executor, which claims its admitted attempt
and reports progress and completion to any daemon replica. The initiating HTTP
request waits at most 30 seconds for launcher admission, not for the command
or its result. An accepted Start is not a readiness guarantee.
| Deployment / action | Start, Stop, Nuke | Restart | Diagnostic shell Logs |
|---|---|---|---|
| Standalone hybrid | Existing local/delegated behavior | Existing behavior | Bounded query |
| Standalone or HA webhook-only | Existing HTTP(S) webhooks | Existing webhook sequence | Shell commands rejected; Logs webhook supported |
| HA shared-local hybrid | Rejected at startup | Rejected | Not an enabled lifecycle profile |
| HA external delegated hybrid | Attempt-scoped external commands; HTTP(S) webhooks also supported | Unavailable: Stop, inspect, then Start | Independently requires executor-response-v1 and an exact replica origin_url |
Missing shell Logs prerequisites does not block lifecycle dispatch or the output reported by lifecycle executors. Configure only capabilities you actually provide. The Environment tab shows the current capability and outcome separately from instance guidance.
Operator configuration (in addition to the normal HA prerequisites):
daemon:
public_url: https://agor.example.com # reachable from executors; may load-balance replicas
deployment:
mode: ha
ha:
execution_topology: external
execution:
unix_user_mode: delegated
managed_envs_execution_mode: hybrid
executor_command_template: /path/to/reviewed-external-launcher
# Assertion about the external Job's TOTAL lifetime, including termination grace.
# Set only after the launcher/control plane enforce this bound.
environment_command_job_deadline_ms: 485000
session_token_expiration_ms: 900000 # at least 545000Do not simply flip the mode on an existing Cloud installation. The trusted
launcher must recognize environment.lifecycle payloads carrying params.attempt,
perform normal tenant/user/branch and Team admission, create the external executor
Job, and exit promptly after that admission. It must not wait for completion,
reserve a synchronous response, retry the command, or expose application ingress.
Other command families, especially diagnostic queries, retain their own handoff.
The executor image must support the same attempt/report protocol as every daemon.
Keep daemon, database, and executor clocks synchronized for the absolute deadlines.
Deadlines and rollout coordination
Admission and claim budgets are shared fixed defaults, not configurable settings. The longer startup allowance applies to managed environments across Agor; it does not change normal agent execution timeouts. Admission acknowledges the external Job, while the claim budget also covers cold-node scheduling and image pulls.
| Bound | Budget |
|---|---|
| Local launcher admission | 30 seconds |
| Executor claim, measured from durable admission | 180 seconds |
| Command, measured from claim | At most 300 seconds, also capped by the admitted absolute deadline |
| Owned process-group cleanup | TERM, then KILL after at most 5 seconds |
| Result delivery grace | 30 seconds beyond the admitted command and cleanup deadline |
| Final result deadline, measured from admission | 515 seconds |
| External Job total bound | Operator-enforced 305–485 seconds, including termination grace |
For Kubernetes, configure command-specific activeDeadlineSeconds plus
terminationGracePeriodSeconds within the asserted total, with backoffLimit: 0
and restartPolicy: Never. The external bound is necessary when an executor
crashes or commands escape its process group. Agor does not implement provider
containment, discover zombie remote resources, or certify their deletion.
Before enabling hybrid, apply PostgreSQL migration
0101_environment_command_discovery and replace the daemon/executor cohort with
compatible versions. The migration changes only the existing recovery SELECT
policy and partial index to include stopping; lifecycle rules remain application
code, using PostgreSQL row locks or SQLite immediate transactions. No SQLite
schema migration is required. Ordinary tenant scope and branch permissions still
apply. The privileged PostgreSQL monitor discovers routing IDs, then switches to
the corresponding tenant for all state transitions.
Coordinate the external launcher and Job deadline before setting the deadline assertion. Drain active attempts before rolling back; retain remote cleanup diagnostics. For branches with attempt metadata, a configuration-only switch to legacy lifecycle handling is not supported: the generic state-write guard deliberately remains closed. Roll back the compatible daemon/executor cohort together instead. Deploy neither a mixed attempt-capability cohort nor a launcher that waits for command results. No particular Git SHA or provider is hard-coded as an eligibility gate.
Script contract and outcomes
For example, a repository variant can call three scripts:
environment:
version: 2
default: remote
variants:
remote:
start: ./scripts/start-remote.sh {{branch.name}}
stop: ./scripts/stop-remote.sh {{branch.name}}
nuke: ./scripts/delete-remote.sh {{branch.name}}
logs: ./scripts/remote-logs.sh {{branch.name}}Each script must finish; do not start a local development server in the executor. Remote resource identity and credentials belong to your scripts/provider. Make Stop safe to repeat. Agor never automatically retries Start, Stop, or Nuke. Start can optionally return the same tiny runtime result used by ordinary managed environments. Emit exactly one stdout line before exiting:
AGOR_ENVIRONMENT_RESULT={"app":"https://preview.example.com","health":"https://preview.example.com/health"}The JSON is the strict {app?, health?} schema described above. Agor parses
control records from stdout only; stderr remains diagnostic output. Invalid or
ambiguous results produce an unknown outcome, not a claim that nothing
happened. The former AGOR_ENVIRONMENT_RESULT_FILE value containing exactly
{"access_urls":[...]} remains accepted at the executor boundary for upgrades,
but new bridges should use the stdout contract. Without a result, the rendered
app and health values are used. Reported links appear in the Environment tab;
the environment pill opens the reported App, falling back to the static app URL.
A successful Start waits for configured health checks; without
one, the command success and unknown health are displayed separately.
An early healthy response never releases a still-running command’s admission.
A successful Stop/Nuke records script success, not independent proof of cleanup.
Only one lifecycle attempt may be active per branch. Duplicate claims cannot run again; stale reports cannot overwrite a newer attempt. Lost claim/result reports expire to unknown, even after the initiating daemon disappears. The UI retains an output tail of at most 32 KiB per attempt, plus three prior attempts, and marks truncation or missing/incomplete output. Do not print secrets: this diagnostic output is readable by authorized branch viewers.
After failed or uncertain cleanup, Retry Stop is allowed once the current attempt has settled or expired. Start requires an explicit warning acceptance bound to the currently displayed failed/uncertain attempt. Starting again can leave additional billable resources running, and your scripts might no longer be able to stop the previous environment. Confirmation is never sticky:
- REST:
POST /branches/<id>/startwith{"confirmation_of":"<current attempt UUID>"}. - MCP:
agor_environment_startwithconfirmationOf, only after operator consent. - CLI:
agor branch env start <branch> --confirm-previous-attempt <UUID>.
Stale confirmation is rejected; refresh and inspect rather than automatically
retrying. Configuration, variant rendering, archiving, and deletion cannot bypass
an active attempt. HA external repository removal with filesystem cleanup: true
is deliberately unavailable: explicitly stop/inspect and use the existing
branch-by-branch cleanup workflow instead of a filesystem-first bulk operation.
Remote development environments
Sometimes Agor can manage the branch but cannot host its development stack. The daemon may run in a locked-down container without Docker, the application may need more CPU or memory than the Agor host has, or the runtime may need access to a cloud-only network. A managed environment does not have to be a local process: its lifecycle commands can bridge to a remote development environment such as a GitHub Codespace, Coder or Gitpod workspace, DevPod target, Kubernetes namespace, or dedicated VM.
Agor does not ship native controllers for all of those providers. Instead, the current primitives support two integration shapes:
- In
hybridmode, a small repository-owned shell bridge calls the provider API or CLI and prints a structured Start result. - In
webhook-onlymode, lifecycle URLs call a trusted external controller. The current webhook contract has the GET-only and authentication limitations described above, so it is not by itself a production-grade provider control plane.
Each command should finish rather than host a foreground server inside the
executor. Remote does not require dynamic URLs: if your deployment has a
predictable address, keep static app and health templates. When the provider
assigns the address at startup, the script can
report it from Start.
Agor’s own repository contains two working references to adapt:
- Railway: a repository-owned script that drives an explicitly bound Railway preview, checks ownership, and reports its URL.
- GitHub Codespaces: a launcher that manages a whole remote workspace rather than just an app deployment.
Vercel, a Kubernetes namespace, or your own VM are candidates for custom commands too. These are possibilities, not tested integrations; define provider-appropriate Stop and Nuke behavior rather than assuming every platform has a resumable server. There is no Agor provider SDK to install: you or your agent adapt a script, and those adaptations still need review.
Provider-neutral lifecycle pattern
A remote bridge should treat Start as reconciliation, not as unconditional creation:
- Render a stable binding from trusted context such as
branch.id, the exact pushedbranch.ref, and a provider repository identity. - Rediscover provider state on every action. Do not rely on a permanent resource name or forwarded URL: remote workspaces can be stopped, renamed, deleted, or recreated outside Agor.
- Validate the provider owner, repository, ref, and binding marker before resuming, stopping, deleting, or reading logs from a resource.
- Create or resume the workspace, wait with a bounded timeout, then report the current access URLs.
- Make Stop and Nuke idempotent. Stop normally retains remote disk; Nuke deletes only the rediscovered and revalidated resource.
Source transfer is a separate concern. If the remote provider clones from a Git host, the branch or commit must be pushed before Start. Do not silently fall back to the default branch, and do not make a lifecycle bridge auto-push under an ambiguous shared credential.
Before sharing a remote variant
- Trust the code you execute. A command that invokes a repo script grants
that script the command’s credentials. Review script changes, not only the
YAML, and
shellQuoterendered string arguments. - Keep secrets out of configuration. YAML, template overrides, rendered commands, logs, and Start results are not secret stores. Commands receive the invoking user’s Global environment variables.
- Choose the owner. On shared branches, agree who supplies provider credentials and pays for resources. A teammate pressing Stop runs with their own credentials, not the original starter’s.
- Scope every action to the branch. Never use wildcard or project-wide teardown in a shared provider project. Clean up the old environment before rendering a different variant.
- Plan for partial success. A timed-out command may already have allocated billable resources; rediscover before retrying and keep a recovery path. Stop may retain billable storage, and public URLs need their own access controls.
Reporting the App and health endpoints
Forwarded hostnames are often assigned only after the remote workspace starts. A shell bridge can publish them by emitting one control line on stdout:
AGOR_ENVIRONMENT_RESULT={"app":"https://preview.example.test","health":"https://preview.example.test/health"}This is the same Runtime URLs from Start contract
described above. The JSON schema is intentionally only {app?, health?}. Either
key may be omitted:
appis the browser destination. It may require the provider’s login.healthis polled by the Agor daemon and should be returned only if that daemon can reach it. If it is omitted, Agor shows Started; health unavailable rather than claiming the environment is healthy.
A static app such as the provider dashboard is useful while Start is still
running. A successful dynamic result replaces that fallback for the current
lifecycle attempt. Stop, Nuke, or another Start clears/reconciles runtime URLs so
Agor does not retain an old forwarded hostname.
Railway reference implementation
Agor’s repository wires its railway-sqlite variant in
.agor.yml to
launcher.mjs,
a dependency-free Node script with adjacent tests and volume-reset logic. The
Railway bootstrap instructions
describe the image, service, domain, and persistent volume.
The preview application in this example is Agor itself. Hosting Agor on Railway and using Agor to manage your app on Railway are different tasks; for your own app, replace the bootstrap and runtime details, not the lifecycle interface.
environment:
variants:
railway:
description: Owned Railway preview
start: >-
node scripts/managed-environments/railway/launcher.mjs start
--repository {{shellQuote repo.github_slug}}
--ref {{shellQuote branch.ref}} --binding {{shellQuote branch.id}}
stop: >-
node scripts/managed-environments/railway/launcher.mjs stop
--repository {{shellQuote repo.github_slug}}
--ref {{shellQuote branch.ref}} --binding {{shellQuote branch.id}}
logs: >-
node scripts/managed-environments/railway/launcher.mjs logs
--repository {{shellQuote repo.github_slug}}
--ref {{shellQuote branch.ref}} --binding {{shellQuote branch.id}}
app: https://railway.comBinding. The reference adopts an explicitly provisioned preview for one
branch and one trusted operator; it does not provision a service per branch.
bindings.json
maps the Agor branch UUID to the exact repository/ref, Railway project,
environment, service, volume, and domain, and the provider-side
AGOR_MANAGED_BRANCH_ID marker must match. The launcher refuses unknown branches
and mismatched ownership, source, volume, or domain state. Do not reuse the
example’s resource IDs or bind two branches to one service/volume.
Credentials. Store these as the invoking user’s Global environment
variables, never in .agor.yml or command arguments:
| Variable | Purpose |
|---|---|
RAILWAY_API_KEY | Environment-scoped Railway project token (RAILWAY_TOKEN is a fallback). |
RAILWAY_AGOR_ADMIN_PASSWORD | Bootstrap password for the preview app; 15+ characters, at most 72 UTF-8 bytes. Not reset for existing accounts. |
RAILWAY_API_TOKEN | Separate workspace token, needed only for the optional destructive volume reset. |
Lifecycle. Start validates the binding, deploys the pushed branch (not the dirty checkout) or reuses an in-flight deployment, waits for readiness, and reports App and health URLs through the standard Start result; the static dashboard link is only the fallback. Logs reads bounded provider diagnostics without starting compute. Stop removes the GitHub push trigger and drains deployments, preserving the service, domain, data, and storage costs; a later infrastructure apply can restore the trigger. The full variant’s Nuke resets the owned volume (erasing its accounts and data) rather than deleting the project, so back up first and inspect rather than repeat after an interruption.
Limits. Its watch mode polls the pushed branch for source updates; dependency and migration changes need Stop/Start. Its cold-start wait can take up to 18 minutes, which fits the 25-minute standalone budget but not the five-minute HA external command budget. Its lifecycle lock is local to one controller. Keep app authentication enabled on the public URL, pin the reference revision you adopt, and retain its safety tests.
GitHub Codespaces reference implementation
Agor’s own repository includes an experimental, opt-in Codespaces example:
- the
codespaces-sqlite.agor.ymlvariant agor-codespace-launcher.mjs, a dependency-free Node bridge over the officialghCLI- the managed Codespaces devcontainer and SQLite bootstrap
The essential variant shape is:
environment:
variants:
codespaces-sqlite:
start: >-
node scripts/managed-environments/codespaces/agor-codespace-launcher.mjs start
--repository {{shellQuote repo.github_slug}}
--ref {{shellQuote branch.ref}}
--binding {{shellQuote branch.id}}
--wait-seconds 1200
--app-port 5000 --health-port 3000 --health-path /health
--port-visibility public
--emit-health public-only
stop: >-
node scripts/managed-environments/codespaces/agor-codespace-launcher.mjs stop
--repository {{shellQuote repo.github_slug}}
--ref {{shellQuote branch.ref}}
--binding {{shellQuote branch.id}}
nuke: >-
node scripts/managed-environments/codespaces/agor-codespace-launcher.mjs nuke
--repository {{shellQuote repo.github_slug}}
--ref {{shellQuote branch.ref}}
--binding {{shellQuote branch.id}}
app: https://github.com/codespacesUse shellQuote for every rendered value embedded in a shell command; it quotes
the value as one POSIX shell argument instead of treating branch text as syntax.
The reference launcher requires Node, an authenticated gh CLI with Codespaces
access, a registered github.com remote, and a pushed branch ref. It derives the
App URL from live Codespaces port metadata. With --emit-health public-only, it
reports health only when GitHub says that forwarded port is public; private
Codespaces URLs remain authenticated App links and are not polled by Agor.
The repository’s experimental variant explicitly passes --port-visibility public, so after the SSH readiness probe succeeds the launcher asks GitHub to
make both 3000 and 5000 public. GitHub may reject that request under an
organization or repository policy, in which case Start fails with a policy
diagnostic instead of pretending the preview is reachable. Public forwarded
ports are available to anyone who knows their URLs, without GitHub
authentication, and GitHub resets them to private when a Codespace
restarts .
Start therefore re-reads and reapplies the requested visibility every time.
This is safe only as an explicit disposable-preview choice, not merely because
the source repository is public. Port 3000 exposes the whole Agor daemon, not
just /health. The Codespaces bootstrap avoids the repository-wide
admin/admin development credential by generating a different mode-0600
password per Codespace. Retrieve it from that Codespace, sign in, and change it
immediately:
cat ~/.agor-managed/bootstrap-admin-passwordThe value never enters .agor.yml, launcher output, dynamic URLs, or the Git
workspace. An older Codespace/SQLite volume created before this behavior still
has its old credential: rebuilding does not reset an existing user. Change that
password before public exposure, or Nuke and recreate the disposable Codespace.
Every invocation of the launcher’s Start action is a reconcile, not an
unconditional Create. This matters after a timed-out/failed Agor Start: press
Start again and the launcher attaches to the still-existing marker-bound
Codespace. It validates owner/repository ID/ref, resumes it if stopped, waits
for SSH and /health, reapplies port visibility, and returns fresh URLs. A
deleted resource is recreated and a duplicate or identity drift fails closed.
Once Agor records the environment as running, its Start control is disabled;
Restart deliberately runs Stop then Start, retaining and resuming the same
Codespace rather than creating a second one. Neither path pulls newer Git
commits into an already-running workspace or rebuilds its devcontainer; source
synchronization remains a separate explicit operation.
The managed devcontainer explicitly installs the
Dev Containers sshd feature.
This is required for the launcher’s authenticated /health check and runtime
logs: the official gh codespace ssh command cannot connect to a custom image
that has no SSH server. If a Codespace was created from an older version of the
devcontainer, rebuild its container or Nuke it and press Play again.
While a Codespace is building, use Cmd/Ctrl+Shift+P → Codespaces: View Creation Log in the browser/VS Code client, or follow the same provider log from another authenticated machine:
gh codespace logs -c <codespace-name> --followgh codespace logs also relies on the Codespaces SSH transport for a custom
devcontainer. Agor’s Logs action therefore returns a bounded snapshot when that
transport is ready, reports it as temporarily unavailable earlier in creation,
and skips all remote log commands for a stopped Codespace so a log read cannot
resume billable compute. Once the Codespace is available, Logs also retrieves the
nested Compose log tail. From a Codespace terminal, you can inspect the stack
directly:
docker compose -p agor-codespaces-sqlite ps
docker compose -p agor-codespaces-sqlite logs --tail=100 --followThe first Play builds Agor’s development image inside the Codespace. Later starts of the same Codespace reuse its Docker layers unless dependency inputs changed. The reference variant gives the cold path a bounded 20 minutes; missing-SSH configuration errors fail immediately rather than consuming that deadline. Standalone Agor lifecycle execution has a fixed 25-minute safety bound, leaving time for launcher setup and result reporting without adding per-variant timeout or retry machinery. Delegated HA jobs retain their shorter operator-configured execution envelope. GitHub Codespaces prebuilds can accelerate the outer devcontainer, but GitHub does not make Docker-in-Docker available during prebuild creation , so a prebuild cannot absorb this nested image build.
This reference is deliberately narrower than a production shared controller: it
uses the controller’s current GitHub actor, a process-local lock/state directory,
and is not a Cloud/HA credential-sponsorship or distributed-leasing design. Never
put a GitHub token in .agor.yml, template overrides, branch data, URLs, or
launcher output. A production integration should use a tenant- and actor-scoped
credential store plus a durable provider binding and lifecycle lease.
Stop preserves the Codespace and its SQLite volume. Before rendering a different variant, Nuke the Codespaces variant (or delete the Codespace in GitHub); after a variant switch Agor intentionally does not retain the old variant’s cleanup command.
See also
- Containerized execution runs the rendered commands inside isolated containers.
- Multiplayer and execution isolation covers how RBAC and execution modes gate who can Start / Render / edit env config.