AI Provider Data Collection: What Actually Gets Transmitted When You Send a Prompt
When your organization sends a prompt to an AI provider, what actually happens to that data? Not the marketing version — the verified, primary-source version. We went through the official documentation, privacy policies, data protection addenda, and developer docs from six providers to answer this question definitively.
Everything below applies to API and enterprise usage only. Consumer-facing products (Claude.ai Free/Pro/Max, ChatGPT web, Gemini app) operate under completely different — and significantly weaker — privacy terms. Anthropic's October 2025 consumer terms update (opt-out training, 5-year retention) does NOT apply to API or Bedrock/Vertex usage. If you're routing traffic through the API, you're in the “commercial” regime covered here.
Summary Ratings
Anthropic Direct API
Endpoint: api.anthropic.com · Commercial API Terms
What Gets Transmitted
Your full HTTP request hits api.anthropic.com. This includes: the entire system prompt, the full conversation history you pass in the messages[] array, all tool definitions, all tool call results, model parameters (temperature, max_tokens, etc.), your API key (in header), and any file content via Files API. Every byte of that request is visible to Anthropic's infrastructure.
Key Facts
- Model training: NONE — API prompts and completions are NOT used for model training. Commercial Terms state: “Anthropic may not train models on Customer Content from Services.” Exception: customers who explicitly opt into the Development Partner Program voluntarily provide data for training.
- Default retention: 30 days — Anthropic automatically deletes prompts and outputs within 30 days of receipt.
- Safety exception: If a conversation is flagged by trust & safety classifiers, it may be retained for up to 2 years and used to improve abuse detection. Safety classifier scores are retained for 7 years regardless. This is NOT opt-outable.
- Zero data retention: Available — Commercial customers (Team, Enterprise, API) can configure ZDR. Under ZDR, inputs/outputs are not stored beyond immediate inference processing. Safety classifier results are still retained. HIPAA BAA available.
- Third-party sharing: Never — Anthropic explicitly states they do not sell API data to third parties.
- Clio system: Anthropic runs a privacy-preserving analytics system called “Clio” on a subset of first-party API traffic (not Bedrock/Vertex) to monitor for misuse patterns. “Trusted organizations with zero retention agreements” are excluded.
Sources: Commercial Terms (June 2025) · Data retention details · Custom data retention controls · Clio research paper
Amazon Bedrock (Claude, Titan, etc.)
Endpoint: bedrock.amazonaws.com · AWS Service Terms
What Gets Transmitted
Your request goes to bedrock.amazonaws.com (not to api.anthropic.com). AWS manages the entire infrastructure. Anthropic (and any other model provider) has zero access to Bedrock logs, customer prompts, or completions. AWS performed a “deep copy” of the model weights into their Model Deployment Accounts — isolated environments the model providers cannot access.
Key Facts
- Prompt storage: NOT stored — “Amazon Bedrock doesn't store or log your prompts and completions.” This is default behavior, not opt-in.
- Model training: NONE — “Amazon Bedrock doesn't use your prompts and completions to train any AWS models and doesn't distribute them to third parties.”
- Model provider access: Fully isolated — AWS uses Model Deployment Accounts where model providers have no access. Anthropic gets zero visibility into Bedrock customer traffic. This is architecturally enforced, not just contractual.
- Fine-tuning: If you use Bedrock's fine-tuning with your own data, that data stays in your AWS account (S3) under your control. It is NOT used to train base models. Fine-tuned models are encrypted with your KMS key.
- Compliance: SOC 1/2/3, ISO 27001/27017/27018/27701, HIPAA-eligible, FedRAMP Moderate (FedRAMP High for some regions), PCI DSS, GDPR, CSA STAR Level 2.
Sources: AWS Docs — Bedrock Data Protection · Bedrock FAQs
OpenAI API (Direct)
Endpoint: api.openai.com · OpenAI API Terms of Service
What Gets Transmitted
Every API request to api.openai.com includes your full messages array (system prompt, conversation history), model and parameters, any files/images attached, your API key identity, and usage metadata. OpenAI processes this in their infrastructure. Unlike Bedrock, there is no architectural isolation from OpenAI — OpenAI can technically see your prompts and completions (though policy prohibits training on them).
Key Facts
- Model training: API customers' data is NOT used for model training by default. “API data is not used to train OpenAI models unless you explicitly opt in.”
- Default retention: 30 days — OpenAI retains API inputs and outputs for up to 30 days for abuse monitoring. After 30 days, they are deleted.
- Zero data retention: Available — Enterprise customers can request ZDR. Under ZDR, data is not stored beyond immediate inference and not accessed by human reviewers at OpenAI. Requires enterprise sales agreement.
- 2025 court order precedent: In May 2025, a U.S. federal court order (NYT v. OpenAI) required OpenAI to preserve consumer and API content pending litigation. The preservation obligation ended September 26, 2025, but the precedent demonstrates that court orders can override ZDR and retention settings at any time.
- Human review: Trust & safety team may review flagged content. Under standard API (non-ZDR) terms, OpenAI employees can access prompts for safety review. ZDR removes this.
- Compliance: SOC 2 Type II, ISO 27001, GDPR (DPA available), CCPA, HIPAA (BAA available for qualifying usage with enterprise agreement).
Sources: OpenAI Usage Policies · Privacy Policy
Azure AI Foundry (Azure OpenAI)
Endpoint: *.openai.azure.com / ai.azure.com · Microsoft DPA
What Gets Transmitted
Requests go to Microsoft's Azure infrastructure, not OpenAI's servers. Microsoft explicitly states: “Azure Direct Models do NOT interact with any services operated by Azure Direct Model providers, for example, OpenAI (e.g. ChatGPT, or the OpenAI API).” Your prompts stay within your Azure tenant geography. Microsoft processes the data, not OpenAI.
Key Facts
- Model training: NONE — “Customer Data, Prompts, and Completions are NOT used to improve Microsoft or third-party products or services without your explicit permission or instruction.”
- OpenAI access: Zero — “Azure Direct Models do NOT interact with any services operated by Azure Direct Model providers, for example, OpenAI.” This is architectural, not just contractual.
- Abuse monitoring storage: Microsoft retains prompts and completions for abuse monitoring in a logically-separated, customer-specific data store. Managed customers can apply to modify or disable abuse monitoring.
- Modified abuse monitoring: Enterprise customers can apply to disable abuse monitoring (e.g. for healthcare applications using PHI). When approved, Azure OpenAI will not store any prompts and completions for that subscription.
- Data location: Processed within customer-specified geography by region. You control regional deployment.
- Stateful features: If using Assistants API, Responses API, or Stored Completions, data is stored at rest in your Azure tenant resource (AES-256 encrypted). You can delete at any time. This persists until you delete it — not auto-expiring.
- Compliance: Microsoft Products and Services DPA, GDPR, HIPAA BAA, SOC 1/2/3, ISO 27001/27018/27701, FedRAMP High (GovCloud), PCI DSS, CSA STAR. Broadest compliance portfolio of the cloud providers.
Source: Microsoft Docs — Azure AI Foundry Data Privacy
Google Vertex AI (Gemini)
Endpoint: *-aiplatform.googleapis.com · Google Cloud Service Terms & CDPA
What Gets Transmitted
Requests go to Google Cloud's Vertex AI infrastructure. Your prompts (text, code, images, audio, video), system instructions, conversation history, grounding context, and configuration parameters are transmitted over HTTPS to Google's servers. Responses include generated output plus metadata (token counts, safety ratings, finish reasons).
Key Facts
- Model training: NONE — Google's Service Specific Terms are explicit: “Google will not use Customer Data to train or fine-tune any AI/ML models without Customer's prior permission or instruction.” This is legally binding through the Cloud Data Processing Addendum (CDPA). Applies to all managed GA and pre-GA models on Vertex AI.
- Default retention: Inputs and outputs are cached in-memory only with a 24-hour TTL for latency reduction. They are not stored at rest. This cache is isolated at the project level.
- Abuse monitoring: When safety classifiers flag suspicious activity, prompts are stored securely for up to 30 days in the customer's selected region. Authorized Google employees may review flagged prompts. This data is NOT used for training. Customers can request an opt-out — if approved, no prompts are stored.
- Zero data retention: Achievable through multiple actions: disable in-memory caching (project-level API call), opt out of abuse monitoring prompt logging (request form), avoid Grounding with Google Search (which has a non-optional 30-day retention), and avoid Grounding with Google Maps (also non-optional 30-day retention).
- Grounding caveat: If you use Grounding with Google Search, prompts, context, and responses are retained for 30 days and this cannot be disabled. Same for Grounding with Google Maps. Use “Web Grounding for Enterprise” if ZDR is required.
- Human review: Only in the context of abuse monitoring. No routine human review of prompts for quality improvement or model training. Access Transparency logs let you see when Google personnel access your content.
- Data residency: Full regional control. Data stored at rest remains in customer-selected location. ML processing occurs within the region where the request is made. Must use regional endpoints (global endpoints do NOT guarantee data residency).
- Customer managed keys: CMEK supported for datasets, models, endpoints, training jobs, and vector search. Also supports Cloud External Key Manager (EKM) for externally managed keys.
- Compliance: FedRAMP High, SOC 1/2/3, ISO 27001/27017/27018/27701, ISO 42001 (AI-specific), HIPAA (BAA available), PCI DSS v4.0.1, HITRUST CSF via Assured Workloads, CSA STAR.
Sources: Vertex AI Data Governance · Vertex AI Zero Data Retention · Vertex AI Abuse Monitoring · Google Cloud Service Terms
Ollama (Self-Hosted / Air-Gapped)
Endpoint: localhost:11434 · MIT License · Your hardware, your rules
What Gets Transmitted
Nothing. When you run Ollama, all inference happens locally on your CPU/GPU. Prompts, responses, and model weights never leave the machine during normal operation. The only outbound network call is a version update check from the desktop app (which sends OS and architecture info) — and this can be disabled with OLLAMA_NO_CLOUD=1. Ollama wraps llama.cpp to perform inference directly on your hardware.
Models are downloaded from registry.ollama.com over HTTPS when you first ollama pull a model. After that, the model runs entirely offline. For air-gapped environments, you can copy the binary and model files via USB from an internet-connected staging machine — Ollama requires zero network access to run inference.
Key Facts
- Data locality: Prompts and responses are processed entirely in local RAM/VRAM. Nothing is transmitted to any external server. Period.
- Telemetry: Minimal. The only outbound call is an auto-update check from the desktop app. Confirmed by maintainer Bruce MacDonald in GitHub Issue #2567: “Ollama does not track any of your data or input. This is the only outgoing call.” Set
OLLAMA_NO_CLOUD=1to disable it entirely. - Air-gapped operation: Fully functional offline after model download. Download binary from GitHub releases + pull models on a connected machine, transfer via sneakernet, run
ollama serve. Zero internet required for inference. - Model storage: Plain binary blobs in a content-addressable store using SHA256 digests at
~/.ollama/models. No encryption at rest — you need OS-level encryption (LUKS, FileVault, BitLocker) if required. - Open source: MIT-licensed, fully auditable. Every network call is visible in the source code.
The Security Reality
Ollama is the ultimate data privacy solution — nothing leaves your machine — but it is NOT a turnkey secure platform. Here is what enterprises need to know:
- No authentication: The API has zero auth. Anyone who can reach port 11434 can pull, push, delete models, and run inference. This was explicitly declined by the Ollama team (Issue #849, closed as NOT_PLANNED). CVE-2025-63389 (CVSS 9.8 Critical) confirms this: all API endpoints lack authentication through at least v0.12.3.
- No TLS: HTTP only on port 11434. No built-in TLS support. You must put a reverse proxy (Nginx, Caddy) in front for encrypted transport.
- Docker defaults are dangerous: The Docker image runs as root and binds to
0.0.0.0by default. Combined with zero auth, any vulnerability becomes RCE-as-root. Cisco Talos research found over 1,000 publicly exposed Ollama instances via Shodan. - Model provenance is weak: No cryptographic model signing. Community models are unverified. Pillar Security demonstrated that attackers can embed malicious instructions in GGUF template metadata that execute during inference, altering outputs invisibly.
- No audit logging: Ollama does not log who requested what, when, or from where. For SOC 2, HIPAA, or FedRAMP environments, you must build logging externally.
- No multi-user isolation: Single-tenant design. No users, no RBAC, no per-model access control, no rate limiting. All callers share the same everything.
- GPU memory not cleared: Ollama does not explicitly zero prompts from GPU VRAM after inference. Known VRAM leak bugs (pre-v0.7.0) can cause data to persist. On shared GPU systems, one user's prompt data could remain accessible.
- Vulnerability track record: Multiple critical CVEs including CVE-2024-37032 (“Probllama”) for RCE via path traversal, CVE-2024-28224 for DNS rebinding that bypasses localhost restriction in under 3 seconds, and several GGUF parser vulnerabilities.
Sources: Ollama GitHub · Wiz — Probllama RCE · Issue #849 — Auth NOT_PLANNED · Issue #2567 — Telemetry confirmation
Side-by-Side Comparison
Enterprise Guidance for Government Workloads
For any government-adjacent or regulated healthcare workload, the hierarchy from strongest to most lenient enterprise data protection is: Ollama (air-gapped) > AWS Bedrock ≥ Azure AI Foundry ≥ Google Vertex AI > Anthropic Direct API ≥ OpenAI API.
Ollama (air-gapped) is the only option where no data leaves your infrastructure at all. For workloads where absolute data sovereignty is non-negotiable — classified environments, SCIF facilities, fully disconnected networks — this is the only path. The trade-off is smaller models, no frontier-model capabilities, and significant security hardening required at every other layer.
AWS Bedrock is the standout cloud option for maximum isolation: model providers literally cannot see your prompts. No storage by default, FedRAMP High available, HIPAA eligible, CMEK, VPC isolation via PrivateLink. Architecturally enforced, not just contractual.
Azure AI Foundry is the strongest option in the Microsoft ecosystem with the broadest compliance portfolio and FedRAMP High via GovCloud. OpenAI has zero access to your Azure prompts.
Google Vertex AI offers a strong combination of contractual protections, regional data residency, and the broadest AI-specific certification (ISO 42001). The grounding retention caveat is worth noting for ZDR-sensitive workloads.
Direct APIs (Anthropic, OpenAI): Appropriate for non-regulated commercial workloads where you want the latest models first. Negotiate ZDR if handling sensitive data. OpenAI's 2025 court-ordered retention is the key risk factor.
The Elephant in the Room: AI Coding Tools
Everything above covers API-level access. But there is a category of tools that most enterprises overlook: AI-powered coding CLIs that see your entire codebase. Claude Code, OpenAI Codex CLI, Gemini CLI, and GitHub Copilot are installed on developer machines and have read access to source files, terminal output, environment variables, and git history. They transmit this context to provider APIs with every request.
Here is what actually happens with your code when you use each tool:
Claude Code (Anthropic CLI)
Hits api.anthropic.com directly (the Anthropic first-party API). Sends all user prompts, file contents Claude reads, terminal command outputs, and full conversation context. Can optionally route through Bedrock, Vertex, or Azure — but the default is direct API.
- Consumer plans (Free/Pro/Max): Anthropic trains on your sessions by default since the October 2025 consumer terms update. Data retained for 5 years when training is opted in. You can opt out at claude.ai/settings/data-privacy-controls, which drops retention to 30 days. Note: even with training opted out, Anthropic still uses content you explicitly rate (thumbs up/down) and safety-flagged conversations for model improvement.
- Commercial plans (Team/Enterprise/API): No training on your code under standard commercial terms. 30-day retention or ZDR with appropriately configured API keys.
- Bug reports submitted via
/bugare retained for 5 years regardless of plan.
Source: Claude Code Data Usage (official)
OpenAI Codex CLI
Supports two auth methods: ChatGPT Sign-In (OAuth through ChatGPT infrastructure) or standard API key (through api.openai.com). Which one you use changes the data policy entirely.
- ChatGPT Plus/Pro (consumer OAuth): OpenAI may use content to train models by default. Opt out via privacy.openai.com.
- API key access: Data is NOT used for training. 30-day retention for abuse monitoring.
- Business/Enterprise: No training. ZDR and Modified Abuse Monitoring available.
Sources: Codex CLI Security · Codex CLI Authentication
Gemini CLI (Google)
This is the most dangerous of the four for unaware users. The default authentication method (“Login with Google” for individual accounts) routes to the consumer Gemini API where data CAN be used for training.
- Individual/unpaid tier: Google states that “prompts, answers, and related code are collected” and may be used to improve Google products and machine learning technologies. Human reviewers may read and annotate this data. The terms explicitly warn: “Do not submit sensitive, confidential, or personal information to the Unpaid Services.”
- Paid Gemini Developer API: Google does NOT use prompts or responses to improve products.
- Vertex AI authentication: No training. Confidential treatment guaranteed.
- No privacy disclosure in CLI: Gemini CLI does not display a privacy notice, data collection disclosure, or opt-out information in its interface. This has been raised as GitHub Issue #1489 regarding GDPR non-compliance.
Sources: Gemini CLI Terms & Privacy · Gemini API Additional Terms
GitHub Copilot
Runs through GitHub's managed Azure OpenAI infrastructure, not the public OpenAI API. GitHub maintains zero data retention agreements with OpenAI, Anthropic, and xAI for model provider access. The most protective defaults of the four tools.
- All plans (including Individual): GitHub states: “By default, GitHub, its affiliates, and third parties will not use your data, including prompts, suggestions, and code snippets, for AI model training.” Individual users can optionally opt in to code snippet sharing for product improvements.
- Business/Enterprise: No training. Prompts and suggestions from IDE completions are not retained. Chat/CLI interactions retained for 28 days. Content exclusion controls available at repo level.
Source: GitHub Copilot Model Hosting
The pattern is clear: every coding CLI that uses a direct provider API exposes your codebase to the provider's data policies. Even with enterprise accounts, the tool is using a direct API key to hit the provider's servers — which means the provider sees your raw source code, prompts, and context. The model provider's contractual promises are the only thing between your proprietary code and their training pipeline.
The Stuff Nobody Documents Clearly
Safety Classifier Retention
Every cloud provider runs automated safety classifiers on 100% of traffic. Anthropic explicitly states classifier scores are retained for 7 years — regardless of ZDR agreements. Azure and OpenAI retain classifier metadata under their respective DPA terms. Google retains abuse monitoring data for up to 30 days. AWS is least transparent about this. Even with ZDR, some metadata about every request persists. The classifiers don't store raw prompt text, but they store derived signals that indicate what type of content was sent.
Flagged Content Gets Much Longer Retention
If your prompt trips a safety/abuse classifier — even a false positive — it exits the normal retention window and can be retained for extended periods (Anthropic: up to 2 years; Azure/OpenAI: DPA-governed; Google: up to 30 days). For enterprise deployments handling sensitive data, this is worth building PII scrubbing into your prompt pipeline regardless of provider.
What Actually Gets Transmitted
When you send any prompt to any cloud provider, the entire HTTP request body is transmitted: system prompt, all message history, tool definitions, function schemas, tool call results, any files, and all model parameters. There is no “partial” transmission — if it's in your API call body, the provider's infrastructure receives it in full. The distinction between providers is what happens to it after receipt. Your defense is prompt-level PII scrubbing before the data ever leaves your infrastructure — or running locally via Ollama where nothing leaves at all.
How Our Platform Addresses This
We built our platform because we saw the same problem every enterprise faces: you want the power of frontier AI models, but every tool on the market asks you to hand your data to providers via direct API keys and hope their policies protect you.
Our platform takes a fundamentally different approach. Here is what we actually do — and why.
No Direct API Keys. Ever. CSP Credentials Only.
This is the core architectural decision that separates our platform from every coding CLI and AI tool on the market. Our platform does not use direct provider API keys to access LLMs. No sk-ant-* Anthropic keys. No sk-* OpenAI keys. No Gemini API keys. Instead, the platform authenticates exclusively through cloud service provider (CSP) credentials — AWS IAM for Bedrock, Azure AD for Azure AI Foundry, Google Cloud IAM for Vertex AI — which provide:
- Stacked authentication: CSP credentials go through your organization's identity provider (IdP) — Okta, Azure AD, Google Workspace — with MFA, session policies, and conditional access enforced at the IdP layer before any model request is made.
- Enterprise model access: By routing through Bedrock, Azure AI Foundry, and Vertex AI, your prompts stay within the enterprise tier where model providers have zero access to your data (architecturally enforced, as documented above). Direct API keys route to provider infrastructure where the provider sees everything.
- Auditable credential chains: CSP credentials produce CloudTrail, Azure Activity, and Cloud Audit logs that your security team controls. Direct API keys produce logs the provider controls.
- No key leakage risk: There are no long-lived API keys to rotate, leak, or embed in code. Credentials are short-lived, scoped, and issued through your IdP.
This is why Claude Code, Codex CLI, and Gemini CLI present a data governance problem: they require direct provider API keys or consumer OAuth, routing your entire codebase through the provider's direct API. Even with enterprise accounts, the provider sees your raw source code.
Code Mode: Our Own Coding Tool
We built Code Mode from scratch specifically to solve the problem documented above. When your developers use Claude Code, Codex CLI, or Gemini CLI, their source code, terminal output, and working context go straight to the provider's direct API. Code Mode routes through our platform instead — which means every coding request goes through DLP scanning, RBAC enforcement, credential isolation, and the audit trail before reaching an LLM via enterprise CSP credentials. Your developers get the same AI-powered coding experience without handing your proprietary code to providers via direct API.
SmartModelRouter: Enterprise-Tier Provider Routing
SmartModelRouter routes every query to the optimal LLM based on task complexity, cost, latency, and compliance requirements. It supports AWS Bedrock, Azure OpenAI, Google Vertex AI, Ollama, and self-hosted models with automatic failover and cost optimization built in. For organizations that require the strongest data protection, SmartModelRouter routes exclusively through enterprise tiers — never through direct provider APIs.
For maximum data sovereignty, SmartModelRouter also routes to Ollama instances running on your own hardware. Sensitive workloads use locally-hosted models where nothing leaves your infrastructure, while less sensitive tasks route to cloud providers for frontier-model capabilities. The routing decision happens per-request, not per-deployment.
DLP Scanning
The platform includes a DLP (Data Loss Prevention) scanner with detection patterns for PII, credentials, API keys, and sensitive data. It scans both inputs and outputs in real time — before prompts leave your infrastructure and before responses are returned to users. This is the “prompt-level PII scrubbing” referenced throughout this report, built into the platform rather than left as an exercise for each team.
Credential Isolation
Provider API keys and credentials are never exposed to agent code. Our platform uses short-lived, scoped credentials injected at runtime with vault-backed secret management. No persistent secrets exist in agent code. Agents make requests through the platform, which injects the appropriate credentials for the routed provider — the agent never sees an API key.
RBAC (Role-Based Access Control)
Access control operates at per-tool, per-model, and per-workflow granularity. OBO (On-Behalf-Of) authentication ensures agents inherit user permissions, so a marketing agent can't access engineering tools and an intern's agent can't use executive-tier models. This is the multi-user isolation layer that Ollama and direct APIs do not provide.
Immutable Audit Trail
Every query, response, tool execution, and approval decision is logged with cryptographic hashing in a tamper-evident audit system. This is the compliance paper trail that cloud providers don't give you: which user sent what prompt to which provider, what model processed it, what DLP patterns were triggered, what the response was, and whether approval gates were satisfied. For SOC 2, HIPAA, and FedRAMP environments, this audit trail is what closes the gap between provider privacy policies and your organization's compliance requirements.
For organizations running government workloads or regulated industries, this is the difference between trusting provider policies and verifying your own controls.
Verified Primary Sources
- Anthropic — API data retention (official)
- Anthropic — Zero data retention agreements
- Anthropic — Custom data retention controls for Enterprise
- Anthropic — Clio: Privacy-preserving analytics on API traffic
- Anthropic — Consumer terms update (confirms API exclusion)
- Anthropic — Commercial Terms of Service (June 2025)
- Anthropic — Consumer Terms of Service (October 2025)
- Claude Code — Data Usage documentation
- AWS Docs — Bedrock Data Protection
- Amazon Bedrock FAQs
- Amazon Model Training & Privacy
- OpenAI Privacy Policy (updated Feb 6, 2026)
- OpenAI Usage Policies
- Microsoft Docs — Azure AI Foundry Data Privacy
- Google Cloud — Vertex AI Data Governance
- Google Cloud — Vertex AI Zero Data Retention
- Google Cloud — Vertex AI Abuse Monitoring
- Google Cloud Service Terms (Section 18: Training Restriction)
- Google Cloud GenAI Privacy Whitepaper (Oct 2024)
- Ollama — GitHub Repository (MIT License)
- Ollama — Telemetry confirmation (Issue #2567)
- Wiz — Probllama: Ollama RCE (CVE-2024-37032)
- Ollama — API Key Authentication (Issue #849, NOT_PLANNED)
- Claude Code — Data Usage (official)
- OpenAI Codex CLI — Security Documentation
- OpenAI Codex CLI — Authentication Methods
- Gemini CLI — Terms of Service and Privacy Notice
- Gemini API — Additional Terms (paid vs unpaid training policies)
- Gemini CLI — No Privacy Notice (GitHub Issue #1489)
- GitHub Copilot — Model Hosting and Data Agreements
This report was compiled February 2026 from verified primary sources. AI provider policies change frequently. Verify against current provider documentation before making legal or compliance decisions.