Skip to main content
Enterprise Data apps need a valid LIGHTDASH_LICENSE_KEY set on your instance before any of the configuration below takes effect. See enterprise features for applying the key.
Data apps are AI-generated React code, built inside an isolated sandbox and stored in an S3-compatible bucket. To enable the feature on a self-hosted instance you need a sandbox provider, a model provider, and a bucket Lightdash can write to. Claude is the default coding agent, and you can run data apps with OpenAI Codex instead. Sandbox providers are configured separately and cover several runtimes — E2B, AWS Lambda MicroVMs, Azure Container Apps, and Google Cloud Run. See sandboxes for the full provider setup; this page covers everything else data apps need.

Prerequisites

  • Enterprise license - LIGHTDASH_LICENSE_KEY must be set on your instance.
  • S3-compatible storage - a bucket Lightdash can write to for app source and built artifacts. If your instance isn’t already using S3, set that up first.
  • A configured sandbox provider - see sandboxes.
  • A coding-agent provider - Anthropic, OpenAI, or an AWS Bedrock account with access to the model used by your selected agent.

Configuration

Add the following environment variables to your Lightdash deployment: Choose one of the coding-agent and provider combinations below.

Claude through Anthropic (default)

Leave APPS_CODING_AGENT unset or set it to claude, then provide an Anthropic API key. Users can choose Sonnet, Opus, or Haiku for each generation. Sonnet is the default. With ANTHROPIC_BASE_URL, Lightdash uses ANTHROPIC_API_KEY as a bearer token and allows the gateway hostname through the sandbox firewall. The gateway must implement Anthropic Messages. See Corporate LLM gateways for authentication and API-path requirements.

Codex through OpenAI

Set APPS_CODING_AGENT=codex and provide an OpenAI API key. Codex uses OpenAI directly whenever AI_DEFAULT_PROVIDER is not bedrock. You do not need an Anthropic API key for Data app generation in this mode. Users can choose GPT-5.6 Sol, Terra, or Luna for each generation. Terra is the default.

Claude or Codex through Bedrock

Set AI_DEFAULT_PROVIDER=bedrock to route the selected coding agent through AWS Bedrock or a Bedrock-compatible gateway. Set APPS_CODING_AGENT to claude or codex, then configure the region and credentials used by the selected route. For Claude, enable the Claude models you want to use in the selected region. For Codex, enable the corresponding OpenAI model IDs, such as openai.gpt-5.6-terra, through the Amazon Bedrock Mantle path. See OpenAI’s Amazon Bedrock guide for supported models and authentication requirements. The Bedrock credentials are the same ones used by AI Analyst - see AWS Bedrock configuration for the full reference. The sandbox firewall automatically allows only the provider endpoints required for the selected agent and region. When BEDROCK_BASE_URL is set, the gateway must support the protocol used by each selected consumer. Claude uses Bedrock’s streaming Invoke API, while Codex uses the OpenAI Responses API through a custom gateway provider. Codex requires BEDROCK_API_KEY in gateway mode and sends model IDs with the openai. prefix, such as openai.gpt-5.6-terra; the gateway must register or translate those names. See the required gateway APIs matrix.
AI_DEFAULT_PROVIDER is an instance-wide setting. Setting it to bedrock also routes AI Analyst through Bedrock. APPS_CODING_AGENT changes only the coding agent used by Data apps.

Gateway networking

The Lightdash backend and the Data app sandbox make separate connections to the gateway. E2B and Azure Sandboxes receive the configured gateway hostname in their dynamic egress allowlist. AWS Lambda MicroVMs use a pre-provisioned egress connector instead, so LAMBDA_MICROVM_EGRESS_CONNECTOR_ARN must permit the gateway hostname. Docker and Cloud Run follow their existing runtime network policy. See LLM gateway egress. Restart the backend. The “Data apps” entry will appear in the New menu for users with the appropriate permission scope.

Optional configuration

Private-network external connections

External connections block private and internal network addresses by default. To let a data app call a trusted internal HTTPS API, set APP_RUNTIME_EXTERNAL_CONNECTION_ALLOWED_PRIVATE_HOST_CIDRS in the Lightdash backend’s deployment environment:
This example allows private addresses in 10.20.0.0/16 for api.internal.example and only 10.30.1.5 for auth.internal.example. Use the hostnames and narrowest address ranges needed by your services. Restart the backend after changing the variable; malformed entries cause a configuration error at startup.
  • Each entry pairs an exact hostname with an IPv4 or IPv6 CIDR. Do not include a URL scheme, port, path, or wildcard. Hostnames are case-insensitive and a trailing DNS dot is ignored.
  • Repeat a hostname to approve multiple ranges, for example api.internal.example@10.20.0.0/16,api.internal.example@fd12:3456::/48.
  • A private destination must match both the hostname and one of its CIDRs. Lightdash checks every DNS answer and rejects the request if any private address is outside the approved ranges. Public addresses keep their existing behavior.
  • For OAuth 2.0 client credentials, approve the token endpoint’s hostname and private address range separately if it differs from the API host.
  • The backend must already have network access and DNS resolution for both services. HTTPS remains required, including a valid certificate trusted by the backend. This setting does not create network routes or disable certificate verification.
After restarting, create the connection under Project Settings → Data app connections with the API’s HTTPS base URL, authentication, and allowed methods and paths. Use the wizard’s test request to verify it before linking an app. See testing a connection.
This is an instance-wide trust policy for external connections across all projects. A hostname entry covers every HTTPS port, and explicitly approved ranges can include sensitive internal services, loopback, or cloud metadata addresses. Only approve destinations you trust with data from linked apps, and keep connection methods and paths appropriately restricted.
Leaving the variable unset or empty preserves the existing private-network block. The exception applies only to requests through external connections, including connection tests and OAuth client-credentials token requests. Connection permissions, credential handling, DNS pinning, redirect blocking, request limits, and rate limits remain in place. Direct browser image loading, sandbox networking, and other backend requests are unaffected. If a test returns blocked_ip, check the exact hostname and every address it resolves to from the backend against the configured CIDRs. For Failed to obtain OAuth access token, also check the token endpoint’s entry, network access, certificate, and client credentials.

Costs

Self-hosting data apps means you pay your sandbox provider and your selected model provider directly:
  • Your sandbox provider bills for sandbox runtime. A typical build runs for 1–15 minutes; sandboxes are paused between iterations and resumed on follow-up prompts.
  • Anthropic, OpenAI, or AWS Bedrock bills per token. Each generation sends the project’s dbt catalog and the user’s prompt to the selected coding agent, plus any attached charts, dashboards, or images.
Both your sandbox provider and your model provider expose usage dashboards. We recommend setting spend limits on both before rolling the feature out to your users.

Permissions

Data apps follow the same space-based permission model as charts and dashboards. The relevant scopes (view:DataApp, create:DataApp, manage:DataApp) are bundled into the default system roles - but on enterprise instances using custom roles, you’ll need to grant them explicitly. See Custom roles for details.

Troubleshooting

The “Data apps” entry doesn’t appear in the New menu. Check that APPS_RUNTIME_ENABLED=true, LIGHTDASH_LICENSE_KEY is set, and the signed-in user has the create:DataApp scope. Builds fail immediately with a sandbox creation error. Check your sandbox provider’s credentials and template configuration — see sandboxes. Builds fail mid-generation with an Anthropic error. For direct Anthropic, check account usage limits and confirm ANTHROPIC_API_KEY is valid. With ANTHROPIC_BASE_URL, confirm the sandbox can reach the gateway and that it accepts bearer authentication at /v1/messages. Codex builds fail with an OpenAI authentication or model error. Confirm APPS_CODING_AGENT=codex. For OpenAI, verify OPENAI_API_KEY; if OPENAI_BASE_URL is set, the gateway must support the Responses API and the model IDs shown in the Data app model picker. For a Bedrock gateway, verify BEDROCK_API_KEY, /responses support, and the openai.-prefixed model ID. Builds fail mid-generation with a Bedrock error. Confirm BEDROCK_REGION is set to a region where the selected agent’s model is available, and that either BEDROCK_API_KEY or the BEDROCK_ACCESS_KEY_ID / BEDROCK_SECRET_ACCESS_KEY pair is valid. If you use IAM credentials, the principal must have permission to invoke the selected model. With BEDROCK_BASE_URL, also confirm the gateway and sandbox allowlist support the selected agent’s API wire. Codex requires an exact OpenAI Bedrock model ID such as openai.gpt-5.6-terra.