AI policy tracker · Git hosting

Who trains AI
on your code?

The AI-training policy of every major Git host — tracked, cited, and updated.

Before you choose where your code lives, it helps to know who reads it. This tracker covers whether the big Git hosting platforms train AI models on your repositories, with a link to each official policy so you can verify for yourself.

The tracker

Status as of August 15, 2026. Policies change — every entry links its official source.

  • Does not train on your code
  • Trains on your code
  • Opt-in consent required
  • Self-hosted — you control it
Provider Parent & jurisdiction Does it train on your code? Source
gitbuild.dev StennMedia (NL)
European Union
Does not train on your code

Code is never used for AI training, model ingestion, or third-party analysis. The prohibition is contractual and technical — there are no data-sharing relationships with AI companies, and a DPA is included on team plans.

Privacy policy
GitHub Microsoft
United States
Trains on your code

Copilot is trained on public repository content under a license grant in the terms of service with no opt-out. Since April 2026, interaction data from Copilot Free and Pro users — including code from private repositories while it is used — trains models by default unless you opt out in settings. Business and Enterprise data is never used for training.

GitHub Terms of Service (§D.4, §J.3)
GitLab GitLab Inc.
United States
Opt-in consent required

GitLab's privacy statement says it will not train language models on your AI inputs without your instruction or prior consent. GitLab Duo does send code and context to third-party model providers to generate responses, but training requires explicit opt-in. On self-managed instances, AI is off by default and nothing leaves your server.

GitLab privacy statement
Bitbucket Atlassian
United States / Australia
Does not train on your code

Atlassian's AI terms bar subcontractors from training on your inputs and outputs, and its LLM providers (OpenAI, Anthropic, Google) operate under zero-retention agreements that never train on customer data. Atlassian may fine-tune open-source models on de-identified aggregated metadata — not your code — subject to your data-contribution settings. Bitbucket code is not used for training by default.

Atlassian AI terms
Azure DevOps Microsoft
United States
Does not train on your code

Microsoft's Azure commitments state that prompts, completions, embeddings, and training data are not used to train foundation models or improve Microsoft products without your permission. Azure DevOps Copilot features inherit that no-training posture.

Microsoft Azure data privacy
AWS CodeCommit Amazon
United States
Does not train on your code

Amazon Q Developer's FAQ states that for Pro-tier users, content is not used to improve services or train foundation models (the free tier is opt-out). CodeCommit itself is plain Git hosting with no model training, and it returned to full general availability in November 2025.

Amazon Q Developer FAQ
Codeberg Codeberg e.V. (non-profit)
Germany / EU
Does not train on your code

Codeberg's members formally resolved that the forge is not and will not use project or user data to train AI. The same July 2026 vote amended its terms to also prohibit hosting projects that mostly consist of AI-generated code, citing copyright concerns.

Codeberg: Protecting our FLOSS commons
Gitea (self-hosted) Gitea community
None — self-hosted
Self-hosted — you control it

Gitea is software you run on your own servers. There are no AI features and no telemetry by default, so repositories never leave your infrastructure and there is nothing for a vendor to train on.

gitea.com
Forgejo (self-hosted) Forgejo contributors
None — self-hosted
Self-hosted — you control it

Forgejo is open-source software you run yourself, with no AI functionality and no telemetry by default. Project maintainers have formally discussed and opposed AI-generated contributions to the project itself.

forgejo.org
SourceHut sr.ht (Drew DeVault)
United States
Does not train on your code

SourceHut has no AI features and does not train on user code. Its founder is outspokenly opposed to LLMs and has described defending the service against aggressive AI-training crawlers.

sourcehut.org

How to check your own host

Four things to look for before you trust a "we never train" claim.

1. Find the data-use clause

Read the section of the terms or privacy policy that covers product development, training, and AI. That is where the real answer lives — not the marketing pages.

2. Look for "train" and "consent"

The operative words are train, fine-tune, and consent. Check whether training is opt-in — you must agree first — or opt-out, meaning you must actively object.

3. Follow the third parties

Many tools send your code to external model providers for inference. Zero-retention agreements mean they do not train on it — but check who else is in the pipeline before you assume.

4. Check the default, not the promise

A policy that promises no training today can change with a quiet terms update. Note the effective date and re-check after major AI announcements.

Why this matters

Code is the highest-quality text dataset that exists. What a host does with yours is a decision worth reading.

Code is training data

Repositories are exactly what model makers want: real, tested, annotated code. Whether your host shares them is a data-policy decision, not an engineering one.

Defaults change

Terms are updated quietly. A "we never train" promise from a few years ago may not survive today — more than one major provider has already changed position.

The answer is a page, not a promise

Read the actual policy. That is why every entry on this page links its source and carries a review date — verify, then decide.

Veelgestelde vragen

The questions developers ask before trusting a Git host with their code.

Does GitHub train AI on your code?

Yes. GitHub trains Copilot on public repository content under a license grant in its terms of service, with no opt-out. Since April 2026, interaction data from Copilot Free and Pro users — including code from private repositories while you are using it — also trains models by default, unless you opt out in your account settings. Business and Enterprise data is never used for training.

Does GitLab train AI on your code?

Not by default. GitLab's privacy statement says it will not train language models on your AI inputs without your instruction or prior consent. GitLab Duo does transmit code and context to third-party model providers to generate responses, but model training requires explicit opt-in. On self-managed instances, AI is off by default and nothing leaves your server.

Which Git hosts never train AI on your code?

gitbuild.dev, Codeberg, SourceHut, Bitbucket, Azure DevOps, and AWS CodeCommit all state in their official policies that they do not train on your code (Atlassian and Amazon with specific carve-outs). Self-hosted Gitea and Forgejo have no telemetry by default, so there is nothing for a vendor to train on.

What changed for GitHub in 2026?

In April 2026, GitHub's terms changed so that interaction data from Copilot Free and Pro users — inputs, outputs, code snippets, and context, including code from private repositories while it is in active use — trains AI models by default. You can opt out in settings; Business and Enterprise data remains excluded.

Does gitbuild.dev train AI on your code?

Never. Your source code is not used for LLM training, model ingestion, or any form of automated analysis beyond what is required to operate the service. The prohibition is contractual and technical — there are no data-sharing relationships with AI companies.

How can I verify these policies?

Every entry in the tracker links to the official policy page it is based on and carries a review date. Policies change, so check the source and the effective date before making a decision.

Pick a host that reads your code less.

Your code should have one job: running your product. gitbuild.dev never uses it for AI training — contractual and technical, on every plan.

Begin met bouwen