GitHub Copilot's policy for AI training: A governance wake-up call (opens in new tab)
GitHub’s April 2026 policy change will make Copilot Free, Pro, and Pro+ interaction data—including code, prompts, outputs, and context—available for AI training by default unless users opt out. The change highlights governance risks for regulated organizations, especially when protections vary by subscription tier or can be altered through policy updates. The post presents GitLab’s no-training commitment, contractual safeguards, and transparency documentation as a stronger model for enterprise AI governance.
What the GitHub Policy Change Means
- Beginning April 24, 2026, GitHub may use Copilot Free, Pro, and Pro+ data for model training by default.
- Covered data includes:
- User inputs and outputs
- Code snippets
- Associated context
- Interaction data
- Users must actively opt out.
- Copilot Business and Enterprise customers remain exempt under existing contracts.
- Data may also be shared with GitHub affiliates, including Microsoft, for AI development.
- Organizations must review license tiers, settings, contracts, and internal AI governance controls.
Why This Matters in Regulated Industries
- Source code can expose:
- Proprietary business logic
- Internal system architecture
- Sensitive data flows
- Financial algorithms and risk models
- Financial institutions may face intellectual-property and model-risk concerns involving trading strategies, underwriting rules, fraud detection, and credit models.
- Frameworks such as Federal Reserve SR 11-7 and DORA require documented oversight of third-party technology and material changes in vendor practices.
- Public-sector environments governed by NIST 800-53 and FISMA may require sensitive code to remain within controlled boundaries.
- Healthcare organizations must consider HIPAA obligations when development tools interact with clinical or patient-adjacent systems.
- Default opt-in training, individual opt-out requirements, and tier-dependent protections create compliance risks.
Requirements for Enterprise AI Vendors
- Contractual certainty: Vendors should clearly and unconditionally define how customer data is handled.
- Auditability: Organizations need documentation about models, training data, subprocessors, retention, and compliance status.
- Independence from vendor incentives: Customer code should not become training data for systems that may benefit competitors.
- Operational flexibility: Regulated customers may require self-hosting, controlled processing boundaries, or clear procedures for vendor changes.
GitLab’s AI Governance Position
- GitLab states that it does not train AI models on customer code at any pricing tier.
- Its AI vendors are contractually prohibited from using GitLab customer inputs or outputs for their own purposes.
- The GitLab AI Transparency Center documents:
- Models powering its features
- Data handling practices
- Subprocessors
- Retention periods
- Feature compliance status
- GitLab emphasizes cloud and model neutrality, supports self-hosted deployments, and addresses vendor changes through its AI Continuity Plan.
- The post argues that these policies reduce vendor-concentration, compliance, and intellectual-property risks.
Closing the Governance Gap
Organizations should ask every AI vendor:
- Is customer data used for model training?
- Who are the model subprocessors?
- What happens if data practices change?
- Can AI processing remain inside the organization’s infrastructure?
- What indemnification applies to AI-generated output?
The post’s recommendation is to favor vendors that provide durable, contractual, and auditable answers rather than relying on defaults, temporary opt-outs, or policies that can change with short notice.