Enterprise AI Analyst and AI agents need a valid
LIGHTDASH_LICENSE_KEY set on your instance before any of the configuration below takes effect. See enterprise features for applying the key.Prerequisites
- Enterprise license —
LIGHTDASH_LICENSE_KEYmust be set on your instance. - A model provider — OpenAI, Anthropic, Azure AI, OpenRouter, or AWS Bedrock. OpenAI and Anthropic are the most tested.
- Enough context budget on the chosen model — AI Analyst sends the project’s dbt catalog with every request.
Enable AI Analyst
Set the main switch and provide credentials for one provider — the minimal setup uses OpenAI, the default provider:AI_DEFAULT_PROVIDER (openai, anthropic, azure, openrouter, or bedrock) and the matching credentials — see model providers below. Optionally, set ASK_AI_BUTTON_ENABLED=true to add an “Ask AI” entry point in the app UI; without it, users reach agents from /ai-agents.
Model providers
Each provider’s exhaustive variable list lives in the environment variables reference; the notes below cover the choices and gotchas per provider.OpenAI (default)
LeaveAI_DEFAULT_PROVIDER unset or set it to openai, then set OPENAI_API_KEY. gpt-6-sol is the default model; override it with OPENAI_MODEL_NAME. All options: OpenAI configuration.
Behind an OpenAI-compatible LLM gateway (LiteLLM, an internal proxy), also set OPENAI_BASE_URL to the gateway URL and OPENAI_MODEL_NAME to a model your gateway exposes. If the gateway doesn’t support streaming (SSE), set OPENAI_SUPPORTS_STREAMING=false. If it enforces Zero Data Retention, set OPENAI_ZERO_DATA_RETENTION=true.
Anthropic
SetAI_DEFAULT_PROVIDER=anthropic and provide ANTHROPIC_API_KEY. claude-sonnet-5 is the default model; override it with ANTHROPIC_MODEL_NAME. All options: Anthropic configuration.
To send Anthropic traffic through a corporate gateway, add ANTHROPIC_BASE_URL. See Corporate LLM gateways for the required URL, authentication, and API paths.
Azure AI
SetAI_DEFAULT_PROVIDER=azure and point Lightdash at your deployment with AZURE_AI_API_KEY, AZURE_AI_ENDPOINT, AZURE_AI_API_VERSION, and AZURE_AI_DEPLOYMENT_NAME. For reasoning-capable deployments (e.g. o3), also set AZURE_AI_DEPLOYMENT_SUPPORTS_REASONING=true. All options: Azure AI configuration.
OpenRouter
SetAI_DEFAULT_PROVIDER=openrouter and provide OPENROUTER_API_KEY; override the default model with OPENROUTER_MODEL_NAME. All options: OpenRouter configuration.
AWS Bedrock
SetAI_DEFAULT_PROVIDER=bedrock and BEDROCK_REGION (required — the AWS region where the target model is available), then authenticate with either BEDROCK_API_KEY or a BEDROCK_ACCESS_KEY_ID/BEDROCK_SECRET_ACCESS_KEY IAM pair — not both. Enable the corresponding models in the selected region before restarting Lightdash. All options: AWS Bedrock configuration.
To send Bedrock traffic through a corporate gateway, add BEDROCK_BASE_URL. The gateway must support the Bedrock APIs used by every enabled Lightdash consumer. See Corporate LLM gateways.
Corporate LLM gateways
Lightdash can route AI Analyst, Autopilot, and Data app model traffic through an HTTP(S) corporate gateway. Use the base URL for the protocol your gateway exposes:OPENAI_BASE_URLfor an OpenAI-compatible gateway.ANTHROPIC_BASE_URLfor an Anthropic Messages gateway.BEDROCK_BASE_URLfor a Bedrock-compatible gateway.
Anthropic-compatible gateway
ConfigureANTHROPIC_BASE_URL before the API’s /v1 segment. Lightdash accepts a trailing /v1 and removes it, but using the unversioned base avoids ambiguity:
ANTHROPIC_BASE_URL is set, Lightdash sends ANTHROPIC_API_KEY as an Authorization: Bearer token. Without a gateway, direct Anthropic requests continue to use the standard x-api-key header.
Lightdash queries /v1/models to determine which models the credential can access. If the gateway does not expose the Models API, set ANTHROPIC_AVAILABLE_MODELS to the exact comma-separated model names the gateway accepts.
An organization-level BYO Anthropic key cannot be combined with an instance-wide ANTHROPIC_BASE_URL. Remove the organization key to use the gateway, or remove the gateway URL to route that organization through its own Anthropic account.
Bedrock-compatible gateway
Configure the gateway together with the existing Bedrock provider settings:BEDROCK_API_KEY, static AWS credentials, or BEDROCK_USE_DEFAULT_CREDENTIALS=true. A custom Bedrock gateway used by Codex requires BEDROCK_API_KEY; Codex cannot combine an overridden endpoint with IAM/SigV4 credentials.
For Claude Data apps, set CLAUDE_CODE_SKIP_BEDROCK_AUTH=true only when the gateway accepts Claude Code requests without AWS authentication. This setting suppresses credentials inside Claude Code; it does not remove the backend provider’s credential requirement.
Required gateway APIs
One base URL can serve multiple consumers only if the gateway implements every wire protocol those consumers use:
A gateway that supports only Bedrock Converse, for example, can serve AI Analyst chat but cannot serve Claude Data app generation. For the coding-agent configuration and model naming rules, see Data apps.
These gateway settings do not reroute AI writeback or the onboarding agent. Autopilot uses the same provider configuration as AI Analyst. Its scheduler must also be able to reach the gateway.
Validate before rollout
- Confirm the Lightdash backend and any scheduler running Autopilot can resolve and connect to the gateway hostname.
- Open AI Analyst and send a short prompt with the configured model. If model discovery fails, verify
/v1/modelsor setANTHROPIC_AVAILABLE_MODELS. - If Data apps are enabled, generate a small app with each coding agent you plan to support. This verifies the sandbox’s separate network path and API wire.
- If Autopilot is enabled, run it on a disposable project and check its recorded provider, model, and outcome. A successful configuration check does not prove the gateway accepts a model request.
- Check gateway access logs for the expected paths in the table above. Redact authorization headers and tokens from logs.
- Restrict direct provider egress only after all enabled consumers succeed through the gateway.
Verified answers
Verified answers use vector embeddings to match new questions to previously validated ones. Enable embeddings withAI_EMBEDDING_ENABLED=true and pick an embedding provider with AI_DEFAULT_EMBEDDING_PROVIDER (openai, bedrock, or azure). The embedding provider can differ from the chat provider — see verified answers for how they’re used at query time.
Autopilot
Autopilot runs scheduled maintenance for each enabled project using the AI SDK runtime. It needs AI Analyst enabled and a provider configured through the instance model settings or organization AI providers and models. It uses the organization’s visible, available default model, then the configured provider default if available, then another available model. Azure uses its configured deployment directly. There is no per-project model override. Every provider with credentials configured is available to AI Analyst and Autopilot, not onlyAI_DEFAULT_PROVIDER. Lightdash only calls a provider when the selected model belongs to it. To keep Autopilot on one provider, set the organization default model in Settings → Ask AI → General to one of that provider’s models. Use a provider’s *_AVAILABLE_MODELS variable to limit which of its models can be selected.
Enable the Autopilot feature on every API and scheduler process:
ai-autopilot if the enable list already contains other flags, and remove it from LIGHTDASH_DISABLE_FEATURE_FLAGS if present. Restart the API and scheduler after changing environment variables. Keep their provider and Autopilot configuration consistent; the scheduler executes scheduled runs and needs access to the selected provider or gateway.
A user with project management permissions can then enable Autopilot from the project’s Home page. Runs act as the user who enabled it and require that user to retain project management permissions. The AI SDK runtime needs no separate service-account token for Autopilot.
Setup and the activity page show the current provider, model, and key source. Enabling Autopilot checks that a model can be configured, without making a model request. Invalid credentials, unavailable deployments, and gateway errors can still fail the first run. Changing organization settings changes the next run’s selection; previous runs retain their recorded model.
Schedule and limits
Choose a schedule per project in the UI.MANAGED_AGENT_SCHEDULE supplies the fallback cron schedule, defaulting to 0 0 * * *. MANAGED_AGENT_SESSION_TIMEOUT_MS defaults to 600000 (10 minutes), and MANAGED_AGENT_MAX_STEPS defaults to 120 model steps per run. The deadline requests cancellation and prevents queued actions from starting; a content write already in progress must finish its bookkeeping. These limits do not set a dollar spending cap. See provider and usage.
Model qualification and cleanup
Autopilot only flags or deletes content on models that are qualified for it. Every other model runs in observe mode. WhenMANAGED_AGENT_VALIDATED_MODELS is unset, these models are qualified for cleanup:
The default models for OpenAI (
gpt-6-sol) and Bedrock (claude-sonnet-4-5) are not qualified, so Autopilot runs in observe mode unless you choose a qualified model. Set the organization default model, or OPENAI_MODEL_NAME, ANTHROPIC_MODEL_NAME or BEDROCK_MODEL_NAME, to a model in the table.
On Anthropic, Autopilot switches to the newest of Claude Opus 5.5, 4.8 and 4.7 your organization may use, unless the organization default is already one of them. On Bedrock it does the same with Claude Opus 5.5 and 4.7. Other providers follow the organization default. Claude Opus 5.5 on Bedrock is not qualified, so a Bedrock organization that can use it runs in observe mode unless you qualify it in MANAGED_AGENT_VALIDATED_MODELS or leave it out of BEDROCK_AVAILABLE_MODELS.
Azure deployments, Google, OpenRouter and any model not listed above start in observe mode. To qualify one, set MANAGED_AGENT_VALIDATED_MODELS to a JSON array of exact provider/model pairs with a mode of observe, flag, or cleanup. Setting the variable replaces the built-in list, so include the built-in models you still want; an empty value qualifies no model. Duplicate qualifications use the most restrictive mode; malformed configuration rejects startup. Qualify a model on disposable representative content before increasing its permissions; use the exact IDs shown in Autopilot, including Azure deployment names or Bedrock inference-profile prefixes.
Autopilot uses the more restrictive of the project’s requested cleanup mode and the model’s qualification. A downgrade appears in setup and the run summary. Observe mode does not mean read-only: chart creation and repair still follow the project’s capability toggles. Disable those capabilities as well when evaluating without content changes.
Upgrading from the hosted runtime
Lightdash 2.262.0 removed the hosted runtime. From that version Autopilot always runs on your configured AI provider.- Nothing to change to keep Autopilot running. Provider credentials come from the shared AI configuration described above. Keep the schedule and timeout settings, and review the step limit and model qualifications.
- Remove the old settings. The hosted-runtime credential, the remote skill list and
MANAGED_AGENT_RUNTIMEare ignored. Lightdash logs a warning at startup for each one that is still set. - No configuration rollback. Setting
MANAGED_AGENT_RUNTIME=anthropic-managedcannot restore the hosted runner. To roll back, redeploy a release before 2.262.0 with the credentials it needs.
MANAGED_AGENT_RUNTIME selects the runner there and must match on the API and scheduler processes.
Costs
Self-hosting AI Analyst means you pay the selected model provider directly. Every user question sends the project’s dbt catalog plus conversation history to the provider, so long conversations against large projects use noticeably more tokens than one-shot prompts. Query results also contribute to usage when data access is enabled. Provider dashboards expose usage — set spend limits before rolling out. To keep prompts smaller:- Limit each agent to the explores and fields it needs.
- Start a new thread when you change subjects so unrelated conversation history is not sent again.
- Keep always-included knowledge documents short.
- Enable compact filter expressions as described below.
Compact filter expressions
Compact filter expressions replace verbose structured filter schemas for AI agent and Model Context Protocol (MCP) query tools with a smaller expression syntax. This reduces the filter-related tool definitions and inputs sent to the model. Total savings vary by model, conversation history, caching, and tool usage; the setting does not reduce the catalog, existing conversation history, or query results. Lightdash Cloud enables compact filter expressions automatically. On a self-hosted deployment, enable the feature on every backend/API replica and on any scheduler or worker process that executes AI agent jobs, then redeploy or recreate those containers:LIGHTDASH_ENABLE_FEATURE_FLAGS already contains other flags, append ai-filter-expressions to its comma-separated list. Remove it from LIGHTDASH_DISABLE_FEATURE_FLAGS if present. Lightdash reads both variables at process startup.
Users do not need to create new agents or threads: the next response uses compact filter expressions, including in an existing thread. After enabling the feature for MCP, reconnect or refresh each client so it reloads the query tool definitions.
Permissions
AI Analyst follows the standard project role model: any user with query access to a project can use AI Analyst there. Restrict access at the project level via roles and groups. To limit AI Analyst to specific projects across the instance, setAI_COPILOT_ALLOWED_PROJECT_UUID to a comma-separated list of project UUIDs.
Troubleshooting
The “Ask AI” button doesn’t appear. ConfirmAI_COPILOT_ENABLED=true, LIGHTDASH_LICENSE_KEY is set, and ASK_AI_BUTTON_ENABLED=true if you want the top-bar entry point. Users without query access to any project also won’t see it.
AI Analyst returns provider authentication errors.
Check the API key for the selected AI_DEFAULT_PROVIDER. For Azure, verify AZURE_AI_ENDPOINT, AZURE_AI_API_VERSION, and AZURE_AI_DEPLOYMENT_NAME all match a deployment your key can call. For Bedrock, confirm the region has the target model enabled for your account.
Multi-step conversations fail against an OpenAI-compatible gateway.
If the gateway enforces Zero Data Retention, set OPENAI_ZERO_DATA_RETENTION=true. If it doesn’t support SSE, set OPENAI_SUPPORTS_STREAMING=false.
Verified answers never match.
Confirm AI_EMBEDDING_ENABLED=true and that the embedding provider credentials are set. Tune AI_VERIFIED_ANSWER_SIMILARITY_THRESHOLD if legitimate matches fall below the default 0.6 similarity cutoff.