LLM Integration

The LLM integration exposes built-in and custom HTTPS model providers to attached VMs through an exe.dev integration hostname. New accounts get a default llm integration attached to auto:all, so VMs can call https://llm.int.exe.xyz/v1/models without storing provider API keys on the VM.

Choose one source for each provider:

  • exe.dev LLM gateway: use exe.dev managed credentials and your exe.dev LLM allocation.
  • API Key: store your provider API key in the integration. The VM can call the integration hostname, but cannot read the key.
  • ChatGPT subscription: connect a ChatGPT account and use it as the OpenAI source for a personal LLM integration. This is not an OpenAI Platform API key.
  • Custom HTTPS endpoint: route supported LLM APIs to another HTTPS provider, with optional bearer-token or custom-header authentication.
  • Disabled: hide that provider from the integration.

ChatGPT subscriptions are only available for the OpenAI provider, and only on personal LLM integrations. Team LLM integrations can use the exe.dev gateway or provider API keys.

Configure in the browser

Open the Integrations page and click LLM.

Integrations page showing the ChatGPT subscription and LLM tiles

Add an LLM integration

  1. Choose an integration name. The default name llm gives attached VMs the hostname llm.int.exe.xyz.
  2. For each provider, choose exe.dev LLM gateway, API Key, ChatGPT subscription, or Disabled.
  3. If you choose ChatGPT subscription for OpenAI, connect a ChatGPT account first, then select the connected account by name.
  4. If you choose API Key, paste the provider key and optionally click Test.
  5. To add a custom provider, choose Add provider → Custom HTTPS endpoint, then enter a unique provider ID, HTTPS base URL, supported APIs, and authentication. Model discovery filters are optional.
  6. Choose where to apply the integration: a VM, a tag, or auto:all.
  7. Click Add integration.
LLM integration modal configured to use a ChatGPT account for OpenAI

The integration hostname is available from any attached VM after the integration is saved.

Connect a ChatGPT account

Connect a ChatGPT account only if you want the OpenAI provider to use a ChatGPT subscription instead of the exe.dev LLM gateway or an OpenAI API key.

Enable device code login in ChatGPT

ChatGPT subscription integrations use ChatGPT's device-code authorization flow. Before connecting an account in exe.dev, make sure device code login is enabled on the ChatGPT side.

Enable device code login in your ChatGPT security settings (personal account) or ChatGPT workspace permissions (workspace admin).

ChatGPT security settings showing device-code authorization for Codex enabled

If this setting is disabled, exe.dev can still show a one-time code, but ChatGPT may reject the authorization before the account is connected.

  1. Confirm device code login is enabled in ChatGPT.
  2. Click ChatGPT subscription.
  3. Enter a short local account name, such as work.
  4. Click Connect account.
  5. Click Open ChatGPT, sign in, and enter the one-time code.
  6. Return to exe.dev and click Done signing in.
ChatGPT subscription modal showing the device code sign-in step

The saved account can now be selected by LLM integrations. You can connect more than one ChatGPT account and choose the account by name.

Configure over SSH

Reinstall the default managed LLM integration:

exe.dev ▶ integrations add llm --name llm --attach auto:all

Create a bring-your-own-key integration. From your local shell, use - to read provider keys from stdin (see Providing secrets):

$ printf '%s' "$OPENAI_API_KEY" | ssh exe.dev integrations add llm --name openai-key --openai=byok --openai-key=- --anthropic=disabled --fireworks=disabled --attach tag:llm

Multiple keys can be supplied together, one per line in the same order as the corresponding flags:

$ printf '%s\n%s\n' "$OPENAI_API_KEY" "$ANTHROPIC_API_KEY" | \
    ssh exe.dev integrations add llm --name team-keys --openai=byok --openai-key=- --anthropic=byok --anthropic-key=-

Add a custom HTTPS provider. --custom-provider-api can be repeated; omit it to enable all supported APIs. Use --header instead of --bearer when the provider needs custom authentication headers.

$ printf '%s' "$CUSTOM_LLM_TOKEN" | \
    ssh exe.dev integrations add llm --name custom-llm \
      --custom-provider=acme=https://api.acme.example/v1 \
      --custom-provider-api=openai_responses --bearer=- --attach tag:llm

Edit an existing integration:

$ printf '%s' "$ANTHROPIC_API_KEY" | ssh exe.dev integrations edit llm --anthropic=byok --anthropic-key=-
exe.dev ▶ integrations edit llm --openai=disabled

To use a ChatGPT subscription for OpenAI, connect a ChatGPT account with the device-code flow. Device code login must already be enabled in ChatGPT before this flow can complete. Enable device code login in your ChatGPT security settings (personal account) or ChatGPT workspace permissions (workspace admin).

exe.dev ▶ integrations setup chatgpt --name work
Open this URL to authorize ChatGPT:
  https://chatgpt.com/...

Enter this code:
  ABCD-EFGH

Waiting for authorization...

After the account is connected, create a personal LLM integration that uses it for OpenAI:

exe.dev ▶ integrations add llm --name chatgpt-llm --openai=chatgpt --openai-account=work --anthropic=disabled --fireworks=disabled --attach auto:all

List, verify, or disconnect ChatGPT accounts:

exe.dev ▶ integrations setup chatgpt --list
exe.dev ▶ integrations setup chatgpt --verify
exe.dev ▶ integrations setup chatgpt --name work --delete

Use from a VM

Attached personal integrations are available at https://<integration-name>.int.exe.xyz. Team integrations use https://<integration-name>.team.exe.xyz.

List available models:

$ curl https://llm.int.exe.xyz/v1/models

Call the OpenAI Responses API:

$ curl https://llm.int.exe.xyz/v1/responses \
    -H "content-type: application/json" \
    -d '{
      "model": "gpt-5.5",
      "input": "Say hello from exe.dev."
    }'

Call the Anthropic Messages API:

$ curl https://llm.int.exe.xyz/v1/messages \
    -H "content-type: application/json" \
    -H "anthropic-version: 2023-06-01" \
    -d '{
      "model": "claude-sonnet-4-6",
      "max_tokens": 256,
      "messages": [{"role": "user", "content": "Hello!"}]
    }'

Transcribe an audio file (up to 25 MB; billed per minute of audio):

$ curl https://llm.int.exe.xyz/v1/audio/transcriptions \
    -F model=gpt-transcribe \
    -F file=@recording.m4a

For Whisper word-level timestamps, request its verbose JSON format:

$ curl https://llm.int.exe.xyz/v1/audio/transcriptions \
    -F model=whisper-1 \
    -F response_format=verbose_json \
    -F 'timestamp_granularities[]=word' \
    -F file=@recording.m4a

Use timestamp_granularities[]=segment for segment timestamps, or repeat the field with both values to receive both. Managed Whisper requests require response_format=verbose_json; its plain JSON response does not include the usage data the gateway needs for billing. Other transcription models use response_format=json. Bare text, srt, and vtt responses and stream=true are rejected for managed requests. Transcription requires managed OpenAI or BYOK; a ChatGPT subscription does not provide it. To keep ChatGPT for chat, attach a separate LLM integration with managed OpenAI or BYOK and send transcription requests to that integration's hostname.

The generic /v1 endpoints route by model ID. To force a provider, prefix the path with /openai, /anthropic, or /fireworks/inference; for example: https://llm.int.exe.xyz/openai/v1/models.

Deepgram transcription

Enable Deepgram in the integration's provider settings, or use integrations edit llm --deepgram=managed over SSH. For your own API key, use integrations edit llm --deepgram=byok --deepgram-key=- and enter the key when prompted.

curl --fail-with-body 'https://llm.int.exe.xyz/deepgram/v1/listen?model=nova-3&smart_format=true' \
  -H 'Content-Type: application/octet-stream' \
  --data-binary @recording.m4a

The transcript is at results.channels[0].alternatives[0].transcript. Uploads are limited to 26 MiB; split larger recordings or submit {"url":"https://example.com/recording.m4a"} with Content-Type: application/json.

As of Sep 2026, Nova-3 recorded audio costs $0.0043/minute per transcribed channel, or $0.0052 with language=multi. Duration includes fractional seconds. See Deepgram pricing for the latest. Managed requests use your exe.dev credits; BYOK requests are billed by Deepgram. Both modes set mip_opt_out=true.

Managed options: model, language, smart_format, punctuate, paragraphs, utterances, utt_split, diarize, multichannel, numerals, profanity_filter, filler_words, encoding, sample_rate, and channels. Managed requests require a known language code or language=multi and reject automatic language detection and paid add-ons. BYOK supports discovered recorded-audio models and their native options. Both modes reject callbacks, live transcription, text-to-speech, and other Deepgram endpoints.

Use with Shelley

Shelley automatically discovers attached LLM integrations through the reflection integration. On new accounts, the default reflection integration is attached to auto:all, so Shelley can see the default llm integration and show its models in the Model: picker without custom model setup or API keys in the VM.

If the model list changes while Shelley is open, choose Add / Remove Models... from the Model: picker and click Refresh. Use Add Model only for separate custom model providers.

If Shelley does not show the integration models, reinstall Reflection with the attached integrations field exposed:

exe.dev ▶ integrations add reflection --name reflection --fields all --attach auto:all

If you use a different LLM integration name, attach that integration to the VM, a tag, or auto:all. Shelley discovers every attached integration of type llm.

Use with Codex

When the OpenAI provider is enabled, run Codex inside an attached VM with the integration as its model provider:

$ codex --model gpt-5.5 \
    -c model_provider=exe-llm \
    -c 'model_providers.exe-llm.name="exe-llm"' \
    -c 'model_providers.exe-llm.base_url="https://llm.int.exe.xyz/v1"'

Or add a provider to ~/.codex/config.toml:

model_provider = "exe-llm"

[model_providers.exe-llm]
name = "exe-llm"
base_url = "https://llm.int.exe.xyz/v1"
requires_openai_auth = false

If your LLM integration uses --openai=chatgpt, the ChatGPT account is connected to exe.dev, not to the VM. Codex still talks to the integration hostname without an OpenAI API key in the VM.

Use with Claude Code

When the Anthropic provider is enabled, Claude Code expects an API key value, so provide a harmless placeholder and point it at the integration hostname:

$ ANTHROPIC_API_KEY=implicit \
    ANTHROPIC_BASE_URL=https://llm.int.exe.xyz \
    claude --model opus

Or add the configuration to ~/.claude/settings.json:

{
  "apiKeyHelper": "printf exe-gateway",
  "env": {
    "ANTHROPIC_BASE_URL": "https://llm.int.exe.xyz"
  }
}

If you use a different integration name, replace llm in the hostname with that name.

Attachments

The default llm integration is attached to auto:all. If you create a separate integration, attach it to a VM, a tag, or all VMs:

exe.dev ▶ integrations attach chatgpt-llm vm:devbox
exe.dev ▶ integrations attach chatgpt-llm tag:llm
exe.dev ▶ integrations attach chatgpt-llm auto:all

See Attaching Integrations for more attachment examples.