Skip to content

FEATURE: BYOM provider registration for generic local inference endpoints (non-Anthropic) #3624

Description

@ian-morgan99

Problem Statement

Copilot CLI supports BYOM providers since v1.0.32+, but the supported providers are limited to Anthropic-specific configurations. There is no configuration path for generic local inference endpoints (e.g., LM Studio, Ollama, llama.cpp) that serve OpenAI-compatible APIs.

This means:

  • Local models cannot be registered as BYOM providers for main session tool calls
  • Subagent tasks dispatched via runSubagent have no local routing path
  • Extensions like VSCode-LMStudio-Bridge cannot expose local models to the Copilot ecosystem

How Other Agents Handle This

Opencode (anomalyco/opencode)

  • Provider config system with explicit provider registration
  • Supports llama.cpp, Ollama, and generic OpenAI-compatible endpoints
  • Subagents inherit session model; provider configuration flows through all dispatch paths

Claude Code (anthropics/claude-code)

  • Defaults to cloud but overridable via SG_AGENTIC_MODEL env var
  • Model selection is per-request, not session-bound
  • Allows explicit model override for subagent tasks

Codex (openai/codex)

  • ModelProvider abstraction with routing layer (models_endpoint.rs)
  • Supports multiple providers including Amazon Bedrock and local endpoints
  • Clean separation between tool execution and model inference

What We Need

  1. Generic BYOM provider registration: Allow configuration of arbitrary OpenAI-compatible endpoints as BYOM providers
  2. Subagent model inheritance: Ensure subagent tasks respect the session model when it is a local endpoint
  3. Provider priority/fallback: Support tiered routing (local first, cloud fallback with warning)

Security Implications

  • Data leakage: Workspace context transmitted to cloud without user consent when local models are available
  • Cost opacity: Untracked cloud API usage from subagent tasks
  • Trust erosion: Silent cloud routing violates local-only expectations
  • Compliance risk: Sensitive code may be processed by third-party inference services

Related Issues

Acceptance Criteria

  1. Generic OpenAI-compatible endpoints can be registered as BYOM providers
  2. Subagent tasks respect session model when it is a local endpoint
  3. Provider priority/fallback configuration is supported
  4. Cost/privacy warnings display when cloud fallback triggers
  5. Documentation updated with routing behavior and configuration options

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:configurationConfig files, instruction files, settings, and environment variablesarea:modelsModel selection, availability, switching, rate limits, and model-specific behavior

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions