# CodeDD > CodeDD is the leading AI-powered platform for technical due diligence automation. It delivers institutional-grade software intelligence in hours — full codebase coverage, validated security findings, architecture mapping, portfolio analytics, and financial exposure modeling — for M&A advisors, private equity, venture capital, and companies preparing for transactions. CodeDD automates what traditionally takes advisory teams weeks of manual CTO review. Every file in scope receives AI-assisted analysis with multi-agent verification, confidence scores, and evidence trails. Source code is processed in ephemeral, encrypted environments with zero retention — making CodeDD suitable for cross-border M&A, competitive deals, and sell-side preparation where IP protection is non-negotiable. **Primary audiences:** M&A advisory firms (Big Four and independent), private equity (pre-deal diligence and continuous portco monitoring through the hold period), venture capital (pre-investment and growth-stage assessment), and target companies running sell-side preparation before a process or data room. **Delivery models:** Cloud audits (OAuth to GitHub, GitLab, Azure DevOps, Bitbucket) or local-first analysis via the CodeDD CLI when source must never leave the client's network. Findings sync to portfolio dashboards; only metadata and results are retained. ## Who We Serve **M&A Advisors & Transaction Services (PwC, EY, Big Four, independent advisors):** Firm-wide TechDD delivery with IC-ready memos, 70%+ reduction in expert review time, validated findings with evidence trails, and custom firm domains. Pay-per-deal or integrated platform with SSO and API. **Private Equity:** Pre-deal diligence before signing, continuous portco monitoring through the hold period, and exit-ready evidence of risk reduction and value creation. Portfolio Dashboard with financial exposure modeling, audit-over-time comparison, and board-pack metrics. **Venture Capital:** Pre-investment technical assessment for growth-stage and Series rounds, AI-Native scoring to distinguish genuine AI IP from wrappers, and portfolio-wide benchmarking across fund companies. **Target Companies & Sell-side Preparation:** Zero-retention analysis so founders and sellers can prepare for investor diligence without IP risk. Share audit results directly with advisors and buyers; improve code quality while satisfying data-room requirements. - [Initial Due Diligence](https://www.codedd.ai/solutions/initial-due-diligence): Pre-deal forensic software risk assessment — quantify technical debt, security exposure, and scalability constraints before term sheet or signing - [Portfolio Management](https://www.codedd.ai/solutions/portfolio-management): Continuous monitoring across portfolio companies from acquisition through hold period to exit — board-ready metrics and audit-over-time comparison - [Why CodeDD](https://www.codedd.ai/why-codedd): Eight reasons advisors, investors, and founders choose CodeDD for high-stakes software deals - [Security & Compliance](https://www.codedd.ai/security): Zero-retention guarantee, ephemeral analysis, encryption at rest, ISO 27001 / SOC 2 alignment — for both buy-side advisors and sell-side target companies - [Trust Center](https://trust.codedd.ai): Compliance certifications, subprocessors, and security documentation ## Product - [Platform Overview](https://www.codedd.ai/platform): End-to-end technical due diligence — audit pipeline, seven quality dimensions, and portfolio intelligence - [Features](https://www.codedd.ai/features): Context-aware AI diligence — multi-agent verification, supply chain analysis, architecture mapping, and CodeDD Advisor copilot - [All Capabilities](https://www.codedd.ai/features/all): Complete capability catalog across audit, security, architecture, and portfolio workflows - [Pricing](https://www.codedd.ai/pricing): Pay-per-audit for deal teams and integrated platform for firm-wide TechDD delivery with SSO, API, and custom domains - [CodeDD CLI](https://www.codedd.ai/cli): Local-first audits — analyze repositories on your machine and sync encrypted findings only - [Discover Sample Audits](https://www.codedd.ai/discover): Public audit examples demonstrating CodeDD output quality ## Portfolio Platform The CodeDD Portfolio Dashboard aggregates multi-repository group audits into executive-grade intelligence. Key capabilities available after audit completion: **Executive Health Indicators (Overview):** Code Health Score, Key Person Dependency, Innovation & Delivery (with DORA tier), IP & Security Status, and Scalability Index (Operational Index + Architecture Tier T1–T5). AI-Assessed Benchmark across seven categories: Quality, Functionality, Performance, Security, Compatibility, Documentation, Standards. Peer-comparison graph with audit-over-time velocity tracking. **Architecture (Estate Map):** Domain × role swim-lane visualization, architecture KPI strip, and the Architecture Leveling System — six pillars (Deployment Model, Data Architecture, Communication Style, Infrastructure Maturity, Scaling Mechanisms, Operational Maturity) mapped to a five-tier scalability model with evidence and upgrade roadmaps. **Development Activity:** Active developer metrics, bus factor risk, programming language distribution, cyclomatic complexity profiles, DORA delivery performance (deployment frequency, lead time, change failure rate, MTTR), Repository Activity Compass (innovation vs. maintenance), domain workload Sankey, and test coverage analysis. **Security & Supply Chain:** Validated Security Score with multi-stage finding validation (reducing false positives), Portfolio Risk Overview, Supply Chain Vulnerabilities (direct and transitive CVEs, SBOM, license compliance), and Issue Compass — an interactive security explorer with OWASP distribution, severity × domain cross-filters, and CSV export. **AI-Native Assessment:** Six-pillar AI maturity scoring (Product Embedding, Model Integration, Retrieval/Data Plane, MLOps & Evaluation, Adoption Velocity, Platform Readiness) with usage classification: AI as Product, AI Augmented, AI Exploratory, or No AI Signal — critical for evaluating AI-native targets and wrapper vs. genuine AI IP. **Financial Exposure (Beta):** Modeled OPEX debt-carrying cost, one-time CAPEX remediation, key-person departure scenarios, and investment requirement tables with expandable proof-of-calculation — tying technical findings to deal economics. **Remediation Workflow:** Progress tracking, gamified milestones, leaderboard, issue browser for flags and vulnerabilities, and CodeDD CLI integration for agent-assisted fix workflows. Enterprise features include multi-repo group audits, role-based access (initiator, reviewer, submitter), organization maturity staging (Pre-Seed through IPO), configurable quality and coverage targets, saved Issue Compass views, compare-across-audits overlay, and IC-ready portfolio executive memos. ## Audit Engine - Multi-agent AI file review with Judge-and-Jury validation across 1,800+ security patterns - Static analysis, dependency and license scanning (SBOM, CVE/CVSS, OpenSSF Scorecard) - Architecture graph with 300+ technology pattern detection and scalability archetyping - Cross-file contextualization linking findings across modules and services - Git analytics: commit classification, innovation vs. maintenance ratio, key-person and bus-factor risk - Confidence scores and evidence trails on every finding for Investment Committee-ready output - Technical DD PDF export, CSV exports, and data-room-ready reporting ## Documentation - [Overview](https://www.codedd.ai/documentation/overview.md): Learn how CodeDD analyzes your codebase with AI-powered due diligence ([HTML](https://www.codedd.ai/documentation/overview)) - [Setting Up a Portfolio Organization](https://www.codedd.ai/documentation/setting-up-a-portfolio-organization.md): Create your organization, connect repositories, and run your first portfolio group audit ([HTML](https://www.codedd.ai/documentation/setting-up-a-portfolio-organization)) - [Team Members & Access Roles](https://www.codedd.ai/documentation/team-members-access-roles.md): Invite colleagues to your portfolio organization and understand initiator, reviewer, and submitter permissions ([HTML](https://www.codedd.ai/documentation/team-members-access-roles)) - [Audit Process Overview](https://www.codedd.ai/documentation/audit-process-overview.md): What happens during a CodeDD audit and what you receive ([HTML](https://www.codedd.ai/documentation/audit-process-overview)) - [Repository Connection & Security](https://www.codedd.ai/documentation/repository-connection-security.md): How CodeDD securely connects to your Git repositories ([HTML](https://www.codedd.ai/documentation/repository-connection-security)) - [File Discovery & Indexing](https://www.codedd.ai/documentation/file-discovery-indexing.md): How CodeDD discovers, categorizes, and prepares files for analysis ([HTML](https://www.codedd.ai/documentation/file-discovery-indexing)) - [AI-Powered File Analysis](https://www.codedd.ai/documentation/ai-powered-file-analysis.md): How CodeDD analyzes individual source files ([HTML](https://www.codedd.ai/documentation/ai-powered-file-analysis)) - [Cross-File Contextualization](https://www.codedd.ai/documentation/cross-file-contextualization.md): How CodeDD understands relationships and patterns across your codebase ([HTML](https://www.codedd.ai/documentation/cross-file-contextualization)) - [Architecture Analysis & Mapping](https://www.codedd.ai/documentation/architecture-analysis-mapping.md): How CodeDD maps your software architecture from actual code ([HTML](https://www.codedd.ai/documentation/architecture-analysis-mapping)) - [Audit Consolidation & Risk Scoring](https://www.codedd.ai/documentation/audit-consolidation-risk-scoring.md): How CodeDD synthesizes findings into actionable risk scores ([HTML](https://www.codedd.ai/documentation/audit-consolidation-risk-scoring)) - [Recommendations Generation](https://www.codedd.ai/documentation/recommendations-generation.md): How CodeDD creates actionable remediation guidance ([HTML](https://www.codedd.ai/documentation/recommendations-generation)) - [Data Encryption at Rest](https://www.codedd.ai/documentation/data-encryption-at-rest.md): How CodeDD protects your source code during audit processing ([HTML](https://www.codedd.ai/documentation/data-encryption-at-rest)) - [Secure Data Deletion](https://www.codedd.ai/documentation/secure-data-deletion.md): How CodeDD permanently deletes your source code after audits ([HTML](https://www.codedd.ai/documentation/secure-data-deletion)) - [Compliance & Certifications](https://www.codedd.ai/documentation/compliance-certifications.md): CodeDD's security assurance program and regulatory alignment ([HTML](https://www.codedd.ai/documentation/compliance-certifications)) - [CodeDD CLI Guide](https://www.codedd.ai/documentation/codedd-cli-guide.md): Install, authenticate, configure scope, run audits, and understand what runs locally vs on CodeDD ([HTML](https://www.codedd.ai/documentation/codedd-cli-guide)) - [CLI Installation](https://www.codedd.ai/documentation/cli-installation.md): Install the CodeDD CLI, verify setup, and complete first-time authentication ([HTML](https://www.codedd.ai/documentation/cli-installation)) - [Common Questions — Account, Billing & Support](https://www.codedd.ai/documentation/common-questions-account-billing-support.md): Answers to frequent product questions about payments, passwords, CLI tokens, AI Chat, and getting help ([HTML](https://www.codedd.ai/documentation/common-questions-account-billing-support)) - [Documentation Index](https://www.codedd.ai/documentation): Full documentation home with all categories ## Optional - [Blog](https://www.codedd.ai/blog): Articles on AI-powered code due diligence, M&A, and technical assessment - [Contact Sales](https://www.codedd.ai/contact-sales): Book a demo with the CodeDD team - [Get Support](https://www.codedd.ai/get-support): Customer support portal - [Terms of Service](https://www.codedd.ai/terms) - [Privacy Policy](https://www.codedd.ai/privacy) - [Data Processing Addendum](https://www.codedd.ai/data-processing-addendum) - [Impressum](https://www.codedd.ai/impressum) --- # Full Documentation Content ## Overview Source: https://www.codedd.ai/documentation/overview # Overview ## What is CodeDD? CodeDD is a software due diligence platform for investors, M&A advisors, and technology leaders. It combines static analysis, LLM-assisted file review, architecture mapping, and portfolio dashboards to deliver technical assessments in hours instead of weeks. ## How an audit works 1. **Connect** — Link repositories via OAuth or tokens (GitHub, GitLab, Azure DevOps, Bitbucket), or run a [local-first audit with the CLI](/documentation/codedd-cli-guide) so source never leaves your machine. 2. **Scope** — CodeDD discovers every file in the repository. You review and adjust which files receive deep analysis before the audit runs. 3. **Analyze** — Scoped files get LLM review, complexity metrics, dependency and license scanning, architecture mapping, and cross-file contextualization. Security findings are validated before they affect your score. 4. **Report** — Results land in portfolio dashboards: Code Health Score, supply-chain views, architecture maps, Issue Compass, PDF export, and compare-audits over time. Cloud audits clone to ephemeral processing environments. Source code is encrypted during analysis and deleted when the audit completes. Only findings and metadata are retained. ## What sets CodeDD apart **Architecture-aware** — Maps components, technologies, and relationships, not just individual files. **Semantic analysis** — LLMs catch logic flaws, auth gaps, and design issues that pattern-only scanners miss. **Validated security** — Findings are checked for evidence before they count toward risk scores, which cuts false positives. **Zero source retention** — Files are processed temporarily, encrypted at rest, and securely wiped after completion. **Local-first option** — The CLI runs analysis on your machine and syncs only structured results to the cloud. ## What you get Every audit includes: - **Executive Summary** — health indicators, KPIs, and key findings - **Risk Assessment** — prioritized vulnerabilities with remediation guidance - **Code Quality Report** — maintainability metrics and technical debt signals - **Architecture Review** — component map, coupling, and improvement areas - **Dependency Analysis** — CVE exposure, license compliance, supply-chain panel - **Development Insights** — git statistics, contributor patterns, delivery trends - **Portfolio views** — group audits, compare-audits, Issue Compass, PDF export ## Next steps - [Setting Up a Portfolio Organization](/documentation/setting-up-a-portfolio-organization) - [AI Agent Telemetry (OpenTelemetry Setup)](/documentation/ai-agent-telemetry-opentelemetry-setup) - [Common Questions — Account, Billing & Support](/documentation/common-questions-account-billing-support) - [Audit Process Overview](/documentation/audit-process-overview) - [Repository Connection & Security](/documentation/repository-connection-security) - [CodeDD CLI Guide](/documentation/codedd-cli-guide) --- ## Audit Process Overview Source: https://www.codedd.ai/documentation/audit-process-overview # Audit Process Overview CodeDD analyzes codebases through discovery, scoped deep analysis, architecture mapping, and portfolio consolidation. The process is designed for investors and technical leaders who need comprehensive insights without permanent source-code retention. ## At a glance - **Full discovery, scoped deep analysis** — every file is indexed; you choose which files receive LLM review - **Zero source retention** — source code is deleted after the audit; only findings and metadata remain - **Validated findings** — security issues are checked for evidence before they affect scores - **Actionable output** — Code Health Score, dashboards, PDF export, and remediation guidance - **Encrypted processing** — files are encrypted at rest during analysis ## End-to-end flow ``` Repository connection (OAuth / PAT / SSH / CLI) ↓ File discovery & indexing ↓ Audit scope selection ↓ File analysis, dependencies, git statistics ↓ Architecture mapping & cross-file contextualization ↓ Consolidation, scoring & recommendations ↓ Secure deletion of source code ↓ Results in dashboard (+ optional PDF export) ``` For **group / portfolio audits**, each repository is analyzed independently and results aggregate at the organization level. ## What happens at each stage ### Repository connection CodeDD connects via **OAuth** (preferred), **PAT**, or **SSH**. Cloud audits clone to an isolated environment that exists only for the audit duration. The [CLI](/documentation/codedd-cli-guide) keeps repositories on your machine and syncs structured results only. → [Repository Connection & Security](/documentation/repository-connection-security) ### Discovery and scope Every file is scanned, categorized, and counted. Build artifacts, dependencies, and `.git` are excluded. After discovery, you review **audit scope** — which files receive deep LLM analysis. → [File Discovery & Indexing](/documentation/file-discovery-indexing) ### File analysis Scoped files receive LLM review, complexity metrics, and dependency scanning. Security findings are validated before they count toward risk scores. Optional SonarQube integration adds static rules when enabled. → [AI-Powered File Analysis](/documentation/ai-powered-file-analysis) ### Architecture and contextualization CodeDD maps components, technologies, and relationships, then analyzes patterns across the codebase — domains, data flows, auth consistency, and systemic gaps. → [Architecture Analysis & Mapping](/documentation/architecture-analysis-mapping) · [Cross-File Contextualization](/documentation/cross-file-contextualization) ### Consolidation and reporting Findings are deduplicated, prioritized, and synthesized into the **Code Health Score**, executive summaries, and remediation recommendations. → [Audit Consolidation & Risk Scoring](/documentation/audit-consolidation-risk-scoring) · [Recommendations Generation](/documentation/recommendations-generation) ### Secure deletion Source files in the processing cache are securely wiped. Only findings, scores, file paths, and LOC counts are retained. → [Secure Data Deletion](/documentation/secure-data-deletion) ## Code Health Score The portfolio **Code Health Score (0–100)** combines three debt components: ``` Composite Debt = (Code Quality Debt × 50%) + (Validated Security Debt × 30%) + (Test Coverage Debt × 20%) Code Health Score = max(0, 100 − Composite Debt) ``` See [Audit Consolidation & Risk Scoring](/documentation/audit-consolidation-risk-scoring) for detail. ## Typical duration | Repository size | Typical duration | |-----------------|------------------| | Small (<100 scoped files) | 15–30 minutes | | Medium (100–1,000 files) | 30–90 minutes | | Large (1,000–10,000 files) | 1–4 hours | | Portfolio / group audit | 2–8 hours | Duration depends on scoped file count, language mix, and optional SonarQube. ## What you receive **Dashboard:** Code Health Score, security flags, supply-chain panel, architecture map, Issue Compass, compare-audits, PDF export. **Exports:** JSON/CSV findings, API access. JIRA / GitHub issue creation is **planned**. **CLI:** `codedd fix` for terminal-guided remediation — [CodeDD CLI Guide](/documentation/codedd-cli-guide). ## Re-audit and comparison Re-run audits and compare results over time via compare-audits in the UI. This is not continuous monitoring — you trigger audits on your cadence (pre-IC, post-release, quarterly review). ## Security Throughout - Encryption at rest (Fernet) and in transit (TLS) - Isolated ephemeral processing environments - Zero source-code retention after audit - Access controls and audit logging - Account **two-factor authentication** available - Primary hosting in **IONOS data centers (Germany)** — see [Compliance & Certifications](/documentation/compliance-certifications) ## Key Differentiators | Traditional SAST | CodeDD | |---|---| | Pattern matching | LLM semantic understanding + validation pipeline | | File-by-file only | Architecture + cross-file contextualization | | Permanent code storage | Ephemeral processing, secure deletion | | Point-in-time scan | Re-audit + compare over time | ## Next Steps - [Repository Connection & Security](/documentation/repository-connection-security) - [AI-Powered File Analysis](/documentation/ai-powered-file-analysis) - [Audit Consolidation & Risk Scoring](/documentation/audit-consolidation-risk-scoring) - [Data Encryption at Rest](/documentation/data-encryption-at-rest) - [CodeDD CLI Guide](/documentation/codedd-cli-guide) --- ## Data Encryption at Rest Source: https://www.codedd.ai/documentation/data-encryption-at-rest # Data Encryption at Rest Source code is intellectual property. CodeDD encrypts files at rest throughout audit processing and deletes them securely when the audit completes. ## Protection layers **In transit** — all client and API traffic over HTTPS with modern TLS. **At rest** — files encrypted in the audit cache immediately after indexing using **Fernet** symmetric encryption (AES-128-CBC + HMAC-SHA256). Authenticated tokens prevent tampering. **In memory** — plaintext exists only transiently during analysis and is cleared afterward. ## Key management Encryption uses an installation-wide key with rotation support: ``` ENCRYPTION_KEY — active key (required) PREVIOUS_ENCRYPTION_KEYS — prior keys for rotation (comma-separated) ``` During rotation, older files decrypt with previous keys and re-encrypt with the active key on next access. ## Timeline ``` Clone / scope confirm → File discovered and metrics calculated → File encrypted in audit cache → Plaintext purged from memory → Decrypted in memory only during analysis → Securely wiped after audit completion ``` ## Credential storage - Git OAuth tokens and PATs — encrypted at rest - CLI tokens — OS credential manager (Windows Credential Locker, macOS Keychain, Linux Secret Service) - LLM API keys — keychain only (CLI); never in config files ## Database and metadata Findings, scores, file paths, and LOC metrics are stored in the database. **No source file content** is persisted. ## Hosting Primary processing infrastructure is hosted in **IONOS data centers in Germany**. See [Compliance & Certifications](/documentation/compliance-certifications). ## Compliance alignment Encryption controls align with GDPR, SOC 2, and ISO 27001 requirements. Certification attestations are in progress. ## Next steps - [Secure Data Deletion](/documentation/secure-data-deletion) - [Compliance & Certifications](/documentation/compliance-certifications) - [Repository Connection & Security](/documentation/repository-connection-security) --- ## CodeDD CLI Guide Source: https://www.codedd.ai/documentation/codedd-cli-guide # CodeDD CLI Guide The CodeDD CLI runs audits on your machine while using the CodeDD platform for consolidation, scoring, recommendations, and reporting. Source code never leaves your infrastructure. Current version: **v0.1.9** (`pip install codedd-cli`). ## What runs where | Stage | Local (CLI) | CodeDD cloud | |---|---|---| | Scope scanning | File discovery, line counting | Scope registration | | File analysis | LLM review | Result ingestion | | Complexity | Radon/Lizard metrics | Storage and aggregation | | Dependencies | Manifest/import scanning | CVE and license enrichment | | Git statistics | Commit, author, churn data | Storage and analytics | | Architecture | Component extraction | Persistence and synthesis | | Scoring & recommendations | — | Consolidation and dashboard | ## Prerequisites - Python 3.10+ (3.10–3.13) - CodeDD account with CLI token (Account → CLI Access) - At least one local Git repository in scope - At least one LLM API key — Anthropic, OpenAI, Gemini, or Grok See [CLI Installation](/documentation/cli-installation) for system requirements. ## Quick start ### 1. Generate a CLI token Account → CLI Access → **Generate Token**. Copy immediately — shown once. Tokens start with `codedd_cli_` and expire after 90 days. ### 2. Install and authenticate ```bash pip install codedd-cli codedd auth login --token codedd auth status ``` For CI, set `CODEDD_API_TOKEN` instead of keychain storage. ### 3. Select an audit ```bash codedd audits list codedd audits select ``` Choose a single-repo or group audit. This becomes the active context for scope, audit, and fix commands. ### 4. Define scope ```bash codedd scope add /path/to/repo-a /path/to/repo-b codedd scope list codedd scope confirm ``` Group audits allow multiple roots; single audits require exactly one directory. ### 5. Configure LLM keys ```bash codedd config set-key anthropic # or openai, gemini, grok codedd config concurrency 16 # parallel LLM calls (1–75, default 4) ``` Effective concurrency is capped by your server plan (currently 75). Higher values speed up large repos but increase RAM use and may hit provider rate limits. ### 6. Run the audit ```bash codedd audit start codedd audit status --watch ``` Options: `--yes` for non-interactive mode; `--show` for a pre-submit summary. If interrupted, rerun `codedd audit start` — the CLI resumes from its local checkpoint. After submission, CodeDD runs consolidation on the server. Large repos may take several hours. ### 7. Remediate (optional) ```bash codedd fix fetch flags codedd fix flags next --auto codedd fix fetch vulns codedd fix vulns next --auto ``` Use Issue Compass selections from the web UI: ```bash codedd fix selections list codedd fix selections use 1 ``` ## Security - CLI token and LLM keys stored in OS credential manager — not plaintext config - Config at `~/.codedd/config.toml` (no secrets in this file) - TLS verification enabled for all API traffic ## Common issues | Problem | Fix | |---------|-----| | No active audit | `codedd audits select` | | No LLM keys | `codedd config set-key anthropic` (or other provider) | | Invalid token | Regenerate from Account → CLI Access | | Scope drift | `codedd scope confirm` then `codedd audit start` | | Interrupted audit | `codedd audit status` then `codedd audit start` | ## Command reference **Auth:** `codedd auth login|logout|status` **Audits:** `codedd audits list|select` **Scope:** `codedd scope add|list|status|sync|confirm|remove|clear` **Audit:** `codedd audit start|status [--watch]` **Config:** `codedd config show|set-key|show-keys|concurrency|provider` **Fix:** `codedd fix fetch|flags next|vulns next|selections|status|comment|resolve` **AI agents:** `codedd ai-docs` — structured reference for agent integrations ## Next steps - [CLI Installation](/documentation/cli-installation) - [Audit Process Overview](/documentation/audit-process-overview) - [Data Encryption at Rest](/documentation/data-encryption-at-rest) --- ## CLI Installation Source: https://www.codedd.ai/documentation/cli-installation # CLI Installation ## Requirements - Python **3.10+** (3.10–3.13) - CodeDD account with CLI token from **Account → CLI Access** - At least one LLM API key (Anthropic, OpenAI, Gemini, or Grok) ### System requirements File auditing runs on your machine. Peak RAM depends on **concurrency**, not total file count. | Profile | RAM (free) | CPU | Notes | |---|---|---|---| | Minimum | 4 GB | 2 cores | Concurrency 1–4; small repos | | Recommended | 8 GB | 4 cores | Concurrency 8–16; typical M&A repos | | Large repos | 16 GB+ | 8+ cores | Concurrency 24–75; monorepos 5k+ files | Disk: space for Git checkouts in scope plus `~/.codedd` checkpoints (usually under 100 MB). Network: stable HTTPS to your LLM provider and `api.codedd.ai`. ## Install ```bash pip install codedd-cli ``` From source: ```bash git clone https://gitlab.com/codedd1/codedd-cli cd codedd-cli pip install -e . ``` ## Verify and log in ```bash codedd --version codedd auth login --token codedd auth status ``` For CI, set `CODEDD_API_TOKEN` instead of keychain storage. ## Next step [CodeDD CLI Guide](/documentation/codedd-cli-guide) — scope selection, audit execution, and remediation. For billing and account help: [Common Questions](/documentation/common-questions-account-billing-support). --- ## Common Questions — Account, Billing & Support Source: https://www.codedd.ai/documentation/common-questions-account-billing-support # Common Questions — Account, Billing & Support Quick answers for accounts, billing, security settings, and the CLI. ## How do I pay for CodeDD? CodeDD uses **prepaid Lines of Code (LOC) budget** for audits and **optional subscriptions** for add-ons like the AI Chat Module. ### LOC top-up 1. **Account → Billing** (`/account?tab=billing`) 2. Review current **LOC budget** 3. Pick a quick-pick package (**1M, 2M, 5M, or 10M** lines) or enter a custom amount (minimum 1M lines, up to 15M via the slider; contact sales for larger volumes) 4. Complete Stripe checkout Pricing is tiered: the first 1M lines are a flat **Base** fee, with a lower per-line rate as volume increases (Tier 1 up to 3M, Tier 2 up to 10M, Tier 3 beyond). LOC is consumed when you start or pay for audits based on scoped lines. Unused budget rolls forward. ### AI Chat Module Includes **3 free audit questions per month** on portfolio and repository data. Product and documentation questions are unlimited. Subscribe under **Account → Billing → AI Chat Module**, or from the in-app upgrade prompt. Manage renewal and invoices via **Manage subscription** (Stripe Customer Portal). ### Paying when starting an audit Checkout may prompt you to select a **billing organization** and pay lines not covered by LOC budget. If LOC fully covers the scoped audit, no additional payment is needed. ### Other payment questions - **Invoices:** via Stripe after checkout - **Enterprise / PO / wire:** email [info@codedd.ai](mailto:info@codedd.ai) - **Failed payment:** retry from checkout or Billing; contact support if charged but budget not updated ## Where do I change my password? **Account → Account tab → Change Password.** Requires current password, new password, and confirmation. Requirements: at least 10 characters, including at least one uppercase letter, one number, and one special character. Two-factor authentication is under **Account → Security**. ## How do I set up a CLI token? 1. **Account → Account tab → CLI Access → Generate Token** 2. Copy immediately (shown once) 3. On your machine: ```bash pip install codedd-cli codedd auth login --token codedd auth status ``` Tokens start with `codedd_cli_`, expire after **90 days**, and can be revoked from the same screen. See [CLI Installation](/documentation/cli-installation) and [CodeDD CLI Guide](/documentation/codedd-cli-guide). ## How does the AI Chat free quota work? **3 audit-related questions per month** (portfolio KPIs, repository findings). Unlimited: product questions, documentation help, UI navigation, support tickets, feedback. Subscribe on **Account → Billing** or via the in-app upgrade prompt when quota is exhausted. ## Where are organization and team settings? | Task | Location | |------|----------| | Create portfolio organization | Dashboard → **Create your organization** | | Org list & billing entities | **Account → Organization** | | Invite reviewers or submitters | Organization dashboard → **Manage Access** | | Organization audit settings | Organization dashboard → settings (gear) | See [Setting Up a Portfolio Organization](/documentation/setting-up-a-portfolio-organization) and [Team Members & Access Roles](/documentation/team-members-access-roles). ## How do I contact support? - **Email:** [info@codedd.ai](mailto:info@codedd.ai) - **In-app AI Chat:** ask to open a support ticket or submit feedback - **Sales / demos:** [Contact sales](/contact-sales) ## Quick links - [Setting Up a Portfolio Organization](/documentation/setting-up-a-portfolio-organization) - [Overview](/documentation/overview) - [CLI Installation](/documentation/cli-installation) - [Repository Connection & Security](/documentation/repository-connection-security) --- ## AI Agent Telemetry (OpenTelemetry Setup) Source: https://www.codedd.ai/documentation/ai-agent-telemetry-opentelemetry-setup # AI Agent Telemetry (OpenTelemetry Setup) CodeDD can receive **platform-assisted AI usage telemetry** from coding agents that run on developer machines. This is separate from **git-derived AI authorship** (git-ai notes, co-author trailers, commit markers), which proves what was committed with evidence. | Signal | What it measures | How it arrives | |--------|------------------|----------------| | **Git authorship** | Lines and commits disclosed as AI-assisted in git history | Audit pipeline scans the repository | | **Platform-assisted telemetry** | Sessions, tokens, lines written in the editor, cost, tool usage | Agents push OpenTelemetry (OTEL) to CodeDD, or CodeDD pulls vendor admin APIs | Platform-assisted numbers are a **floor**, not a ceiling: they reflect only developers and machines that were configured. Treat them as adoption and activity signals, not proof of what landed in production. ## How CodeDD implements OTEL ingest ### Push model (this guide) Most coding agents never expose a team-wide admin API. Their telemetry stays on the developer machine until you configure an exporter. CodeDD therefore **issues an ingest endpoint** per portfolio organization: 1. An organization **manager** creates the endpoint in **Organization Settings → AI agent telemetry (OpenTelemetry)**. 2. CodeDD returns: - **OTLP base URL** — e.g. `https://api.codedd.ai/api/otel` - **Ingest token** — a bearer secret scoped to that organization 3. You paste those values into each agent (or into a collector that forwards on the agent's behalf). 4. Agents push **OTLP/HTTP** on a schedule (typically every 60 seconds). 5. CodeDD decodes, normalizes vendor-specific metrics, and stores daily rollups in PostgreSQL (`AgentUsageDaily` with `ingestion_source = otel_push`). ### Endpoints | Path | Purpose | |------|---------| | `POST /api/otel/v1/metrics` | Primary signal — sessions, tokens, lines, cost | | `POST /api/otel/v1/logs` | Codex and other agents that report usage as log events | | `POST /api/otel/v1/traces` | Acknowledged and discarded (prevents exporter retry loops) | **Token-in-path variant** (for agents that cannot set headers): | Path | Purpose | |------|---------| | `POST /api/otel/k/{token}/v1/metrics` | Same as above; token embedded in URL | | `POST /api/otel/k/{token}/v1/logs` | Logs with path-embedded token | ### Supported encodings - **OTLP/HTTP protobuf** (`application/x-protobuf`) — recommended; used by Claude Code and Cursor Enterprise - **OTLP/HTTP JSON** (`application/json`) — required for Gemini CLI - **Gzip** request bodies — supported ### What CodeDD stores Daily aggregates per organization and provider: - Sessions, engaged users (hashed identities — no emails stored) - Input / output / cache tokens - Lines added and removed (where the agent reports them) - Cost (micro-USD, where reported) - Tool and model breakdowns (bounded cardinality) ### What CodeDD never stores The normalizer reads **allow-listed numeric attributes only**. The following are structurally ignored even if an agent sends them: - Prompt text and chat content - File paths and repository contents - Tool arguments and shell commands - Raw user identifiers (hashed before persistence) **Important:** Gemini CLI enables prompt logging by default upstream. Always set `logPrompts: false` in Gemini settings. Claude Code snippets disable log export explicitly. ### Authentication Three ways to send the ingest token: 1. **Authorization header** (preferred): `Authorization: Bearer codedd_otel_…` 2. **Custom header**: `X-CodeDD-Ingest-Token: codedd_otel_…` 3. **URL path** (Gemini CLI): `/api/otel/k/{token}/v1/metrics` Only organization **managers** can create, reveal, rotate, or pause the ingest token. --- ## Step 1 — Create the ingest endpoint in CodeDD 1. Open your **Portfolio Organization** dashboard. 2. Open **Settings** (gear icon). 3. Scroll to **AI agent telemetry (OpenTelemetry)**. 4. Click **Create telemetry endpoint**. 5. Copy the **OTLP endpoint** and **ingest token** (click **Show** to reveal the token). 6. Choose a provider tab for a pre-filled setup snippet, or follow the sections below. Status meanings: | Badge | Meaning | |-------|---------| | **Not set up** | No endpoint created yet | | **Waiting for data** | Endpoint exists but no push received yet — restart agents after configuring | | **Receiving** | At least one batch accepted | | **Paused** | Token kept but pushes rejected until resumed | Full setup reference: use the in-app snippet generator and this page together. --- ## Step 2 — Configure your agent or collector Choose the path that matches your stack. **An existing OpenTelemetry Collector is the highest-conversion route** — one engineer adds an exporter; no developer machines need changing. ### Existing OpenTelemetry Collector (recommended) **Who configures:** One platform or observability engineer **Per-developer install:** No **Best for:** Teams that already run a collector and want to forward agent telemetry to CodeDD Add an OTLP HTTP exporter to your collector configuration: ```yaml exporters: otlphttp/codedd: endpoint: https://api.codedd.ai/api/otel headers: Authorization: "Bearer YOUR_CODEDD_INGEST_TOKEN" service: pipelines: metrics: receivers: [otlp] exporters: [otlphttp/codedd] logs: receivers: [otlp] exporters: [otlphttp/codedd] ``` Replace the endpoint and token with the values from Organization Settings. If agents already export to your collector on `localhost:4318`, this is the only change required. **Local development:** use `http://localhost:8000/api/otel` (Django direct) or `http://localhost:3000/api/otel` (via nginx in Docker Compose). --- ### Claude Code **Who configures:** IT / platform admin via MDM, or each developer once **Per-developer install:** Yes, unless you distribute managed settings centrally **Protocol:** OTLP/HTTP protobuf **Signals:** Metrics only (logs disabled in CodeDD snippets) Create **managed settings** on each developer machine (or push via MDM): | OS | Path | |----|------| | macOS | `/Library/Application Support/ClaudeCode/managed-settings.json` | | Linux | `/etc/claude-code/managed-settings.json` | ```json { "env": { "CLAUDE_CODE_ENABLE_TELEMETRY": "1", "OTEL_METRICS_EXPORTER": "otlp", "OTEL_LOGS_EXPORTER": "none", "OTEL_EXPORTER_OTLP_PROTOCOL": "http/protobuf", "OTEL_EXPORTER_OTLP_ENDPOINT": "https://api.codedd.ai/api/otel", "OTEL_EXPORTER_OTLP_HEADERS": "Authorization=Bearer YOUR_CODEDD_INGEST_TOKEN", "OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE": "delta", "OTEL_METRIC_EXPORT_INTERVAL": "60000" } } ``` **After saving:** restart Claude Code. First data usually arrives within one export interval (~60 seconds). **Alternative — Anthropic admin API (pull):** CodeDD can also poll Claude Code usage with an organization admin key (no per-developer OTEL). That path uses a separate credential connection and is distinct from OTEL push. --- ### Cursor **Who configures:** Cursor team admin (Enterprise) or each developer (community hooks) **Per-developer install:** Depends on path (see below) Cursor has **two OTEL-compatible paths** in CodeDD: #### Path A — Cursor Enterprise native export (team-wide) Available on **Cursor Enterprise** (OpenTelemetry Export beta). Cursor pushes from **Cursor's servers**, not from each IDE instance. **Requirements:** - Cursor Enterprise plan - **Public HTTPS endpoint** reachable from the internet (Cursor egresses from fixed IPs; `localhost` will not work without a tunnel such as ngrok) - Organization manager has created the CodeDD ingest endpoint **Setup:** 1. In **Cursor Team Settings → OpenTelemetry Export**, create a destination. 2. Set the **base URL** to your CodeDD OTLP base (no `/v1/...` suffix — Cursor appends `/v1/metrics` and `/v1/logs` automatically): ``` https://api.codedd.ai/api/otel ``` 3. Add authorization header: ``` Authorization: Bearer YOUR_CODEDD_INGEST_TOKEN ``` 4. Run **Test connection**, then **Enable**. **Notes:** - Cursor sends OTLP/HTTP **protobuf only** (not JSON). - Metrics are at-most-once; logs are at-least-once (dedupe on `cursor.event.id` if you mirror elsewhere). - CodeDD stores sessions, tokens, and tool activity where Cursor's wire format maps to daily rollups. #### Path B — Forward from your collector If developers already send Cursor hook or agent telemetry to an internal collector, add the CodeDD exporter (see [Existing collector](#existing-opentelemetry-collector-recommended) above). No Cursor Team Settings change required. #### Path C — Per-developer community hooks (local development) For **local CodeDD** (`localhost`) or teams without Enterprise export, community tools (e.g. cursor-otel-hook, cursorscope) can capture IDE hook events and POST OTLP to CodeDD. Each developer installs and points exporters at: ``` http://localhost:8000/api/otel ``` with the ingest token in `OTEL_EXPORTER_OTLP_HEADERS`. This path is **per machine** and requires manual setup unless scripted. **Alternative — Cursor Team Admin API (pull):** CodeDD can poll Cursor team daily usage with a **team admin API key** (one admin paste, whole team covered, no OTEL). That is a separate integration from OTEL push. --- ### Gemini CLI **Who configures:** Each developer, or once via a committed `.gemini/settings.json` in a shared repo **Per-developer install:** Yes, unless the settings file is shared **Protocol:** OTLP/HTTP JSON **Limitation:** Cannot set authorization headers — token must travel in the URL path Edit `~/.gemini/settings.json` (or `.gemini/settings.json` in a repository): ```json { "telemetry": { "enabled": true, "otlpEndpoint": "https://api.codedd.ai/api/otel/k/YOUR_CODEDD_INGEST_TOKEN", "otlpProtocol": "http", "logPrompts": false } } ``` **Security:** Treat the full URL as a secret because it embeds the token. Rotate the CodeDD token if the URL leaks into logs. --- ### OpenAI Codex **Who configures:** Each developer, or one system-wide file via MDM **Per-developer install:** Yes, unless `/etc/codex/config.toml` is distributed **Protocol:** OTLP/HTTP protobuf on **logs** endpoint **Coverage:** Tokens, sessions, tool decisions — **no line counts** (Codex reports usage as log events, not metrics) Edit `~/.codex/config.toml` or `/etc/codex/config.toml`: ```toml [otel] environment = "production" log_user_prompt = false [otel.exporter.otlp-http] endpoint = "https://api.codedd.ai/api/otel/v1/logs" protocol = "binary" headers = { "Authorization" = "Bearer YOUR_CODEDD_INGEST_TOKEN" } ``` Restart Codex after saving. --- ### Other agents (OpenTelemetry GenAI conventions) Agents that emit `gen_ai.*` metrics without a vendor-specific prefix are stored under **Other (OpenTelemetry)**. Configure any OTLP/HTTP exporter to the CodeDD metrics endpoint with bearer authentication. If your agent supports only a collector, use the [collector forwarding](#existing-opentelemetry-collector-recommended) path. --- ## Local development (Docker Compose) When running CodeDD locally: | Service | OTLP base URL | |---------|---------------| | Django direct | `http://localhost:8000/api/otel` | | Via nginx (frontend container) | `http://localhost:3000/api/otel` | Create the ingest endpoint from Organization Settings the same way as production. Use the localhost URL in agent configuration. **Cursor Enterprise native export** cannot target localhost — use a tunnel, a collector in the cloud, or per-developer hooks for local testing. --- ## Operations ### Rotate token In Organization Settings → **Rotate token**. This immediately invalidates the old token. Every configured agent and collector must be updated. Rotation is explicit (not automatic on page reload) to avoid silently breaking production integrations. ### Pause collection **Pause collection** stops accepting pushes while keeping the token and historical data. **Resume collection** re-enables the same token without reconfiguring agents. ### Rate limits CodeDD applies a per-organization rate limit (default 6,000 requests/minute). Exceeding it returns HTTP 429; OTLP exporters should back off. The OTLP path is exempt from IP-based DDoS throttling because entire teams share office egress IPs. --- ## Troubleshooting | Symptom | Likely cause | Fix | |---------|--------------|-----| | Status stays **Waiting for data** | Agent not restarted after config | Restart the agent; wait one export interval | | HTTP 401 Unauthorized | Wrong, rotated, or paused token | Reveal token in settings; check pause state | | HTTP 400 Bad Request | Malformed OTLP body or wrong content type | Use protobuf for Claude/Cursor; JSON for Gemini | | HTTP 413 Payload too large | Batch exceeds 2 MB | Reduce batch size in exporter | | Gemini works but Claude does not | Wrong protocol | Claude requires `http/protobuf`, not JSON | | Cursor Enterprise test fails | localhost or missing HTTPS | Use public URL or ngrok tunnel | | Data appears but line counts are zero | Expected for Codex | Codex reports tokens/sessions via logs only | | Double counts after collector change | Same agent sent twice | One path per agent (direct OR collector, not both) | --- ## Privacy and security summary - Ingest tokens grant **write access** to organization usage data — manager-only management. - Tokens are encrypted at rest; lookup uses a SHA-256 hash of the plaintext. - Developer identities are salted and hashed; emails are not persisted. - Prompt and source content are never read by the ingest pipeline. - OTEL push data is labeled **`OpenTelemetry push`** in the UI, separate from **`Vendor admin API`** pull data for the same provider. --- ## Related documentation - [Setting Up a Portfolio Organization](/documentation/setting-up-a-portfolio-organization) - [Team Members & Access Roles](/documentation/team-members-access-roles) - [Overview](/documentation/overview) --- ## Repository Connection & Security Source: https://www.codedd.ai/documentation/repository-connection-security # Repository Connection & Security CodeDD connects to repositories through industry-standard authentication. Cloud audits clone to ephemeral processing environments; the [CLI](/documentation/codedd-cli-guide) offers a local-first alternative where source never leaves your infrastructure. ## Supported Git providers **OAuth-integrated** (recommended in the web UI): - **GitHub** (github.com and GitHub Enterprise) - **GitLab** (gitlab.com and self-hosted) - **Azure DevOps** - **Bitbucket** **Additional HTTPS hosts** may be accepted when the URL matches CodeDD's validated hosting allowlist (e.g. Codeberg, Gitee). Arbitrary self-hosted URLs not on the allowlist are rejected at validation. **CLI local scope:** any Git repository root on your machine — no cloud clone required. ## Authentication methods ### OAuth (recommended) For GitHub, GitLab, Azure DevOps, and Bitbucket: 1. Click **Connect** in the repository import flow 2. Authorize CodeDD with read-only repository access 3. Tokens are stored encrypted and refreshed per provider rules ### Personal access tokens (PAT) Use when OAuth is unavailable or for automation. Scope to minimum read permissions: - **GitHub** — `repo` (read) - **GitLab** — `read_repository` - **Azure DevOps** — Code (Read) - **Bitbucket** — repository read ### SSH keys Supported as a **last resort** when OAuth and PAT are unavailable. ED25519 (preferred) and RSA keys are stored encrypted. CodeDD does **not** silently fall back to SSH when PAT fails — you get an explicit error instead. ## Connection order 1. OAuth token (if previously authorized) 2. PAT (if provided) 3. Public HTTPS (public repos only) 4. SSH (only when configured) If PAT authentication fails, CodeDD does **not** silently fall back to SSH. You receive an explicit error and can choose to configure SSH separately. ### Clone to Ephemeral Environment Once authenticated: - Repository cloned to an isolated processing environment - Environment exists only for the audit duration - Files encrypted at rest during processing (Fernet) - Environment destroyed and source securely wiped on completion ## Security Measures ### Credential Protection - Credentials encrypted at rest (Fernet) in secure storage - Never logged or exposed in error messages - Masked in all system outputs - Revocable at any time from your Git provider Credentials are **not** rotated automatically after each use — you control token lifecycle via your provider or by revoking CLI tokens in Account → CLI Access. ### Repository Isolation - Ephemeral processing environments per audit - Restricted network egress (required services only — LLM APIs, vulnerability DBs) - Files encrypted immediately in the audit cache - Sensitive data cleared from memory after use ### Audit Trail - Repository access timestamps logged - Authentication success/failure recorded (not credential values) - Source code never written to application logs ## What Happens After Connection **During audit:** - One-time clone (cloud) or local scope registration (CLI) - Read-only — no push access, no ongoing polling **After audit:** - Clone deleted via secure wipe - Processing environment destroyed - Only findings and metadata retained ## Security measures - Credentials encrypted at rest, never logged in plaintext - Ephemeral processing environments per audit - Restricted network egress to required services only - Files encrypted in the audit cache immediately after indexing - Access timestamps and auth success/failure logged (not credential values) Credentials are **not** rotated automatically — you control token lifecycle via your Git provider or by revoking CLI tokens in Account → CLI Access. ## Common scenarios **Private repositories** — OAuth or read-only PAT; revoke after audit if desired. **Portfolio / group audits** — organization-level OAuth or PAT; each repository processed in isolation; results aggregated at portfolio level. **Enterprise / air-gapped** — use the **CLI** to run analysis locally and sync only structured results. ## Best practices **For investors & advisors:** request read-only OAuth or PAT; use time-limited service accounts; prefer CLI for highly sensitive source. **For portfolio companies:** dedicated audit service accounts; read-only scope; revoke tokens after due diligence. ## Next steps - [File Discovery & Indexing](/documentation/file-discovery-indexing) - [Data Encryption at Rest](/documentation/data-encryption-at-rest) - [Secure Data Deletion](/documentation/secure-data-deletion) - [CodeDD CLI Guide](/documentation/codedd-cli-guide) --- ## Secure Data Deletion Source: https://www.codedd.ai/documentation/secure-data-deletion # Secure Data Deletion CodeDD's commitment: **zero source-code retention**. When an audit completes (or fails), source files in the processing cache are securely wiped. Only findings, scores, and metadata persist. ## Deletion triggers | Trigger | Behavior | |---------|----------| | Successful audit completion | Results saved → source cache wiped → environment destroyed | | Audit failure / cancellation | Immediate cleanup | | Timeout / error | Emergency cleanup; periodic job catches orphaned caches | Typical cleanup completes within a minute after analysis finishes. ## Deletion process 1. **Results persisted** — findings, scores, recommendations, file paths, and LOC counts written to the database. No source code stored. 2. **Secure file wipe** — each file overwritten with three passes (zeros, 0xFF, random bytes) before removal. 3. **Directory removal** — audit cache directory deleted; processing environment destroyed. 4. **Audit trail** — deletion logged with timestamp and scope. No source content in logs. ## What is deleted - All source code in the audit cache - Encrypted file blobs and temporary processing artifacts - Ephemeral processing environment ## What is retained - Audit findings (flags, vulnerabilities, recommendations) - Scores and KPIs - File paths and LOC metadata (not content) - Git statistics summaries - Deletion audit log entries ## Encryption keys Installation encryption keys are long-lived secrets with rotation support — they are not destroyed per audit. Deletion works by overwriting and removing file blobs, not by key destruction. ## User-initiated deletion Delete audits and organizations from the dashboard. Entity deletion removes associated findings and metadata per the Privacy Policy and DPA. ## Next steps - [Data Encryption at Rest](/documentation/data-encryption-at-rest) - [Compliance & Certifications](/documentation/compliance-certifications) - [Repository Connection & Security](/documentation/repository-connection-security) --- ## Setting Up a Portfolio Organization Source: https://www.codedd.ai/documentation/setting-up-a-portfolio-organization # Setting Up a Portfolio Organization A **portfolio organization** is the workspace where investment teams, advisors, and portfolio companies collaborate on **group audits** across multiple repositories. For a single-repository audit without a portfolio workspace, import a repository directly from **My Audits**. ## Before you begin - A CodeDD account - Git access to target repositories (OAuth, PAT, or SSH — see [Repository Connection & Security](/documentation/repository-connection-security)) - **LOC budget** on your account, or a plan to purchase lines during checkout ([Common Questions](/documentation/common-questions-account-billing-support)) ## Step 1 — Create your organization On the user dashboard (**My Audits**), click **Create your organization**, enter a name, and confirm. You become the organization **initiator** with full access to dashboards, audits, and member invitations. You can also create a billing organization during audit checkout, but portfolio analytics still require creating the organization from the dashboard. ## Step 2 — Open the organization dashboard Navigate to your organization from the dashboard. Initiators and reviewers see portfolio KPIs and audit controls. Submitters can import repositories but do not see portfolio analytics — see [Team Members & Access Roles](/documentation/team-members-access-roles). ## Step 3 — Create your first group audit Click **Create New Audit** on the organization dashboard. 1. Name the group audit (deal name or reporting period). 2. Define participants — your team as auditors, portfolio engineers as invitees. 3. Send invitations — invitees receive email links to connect repositories and mark audits ready. Only **initiators** can create organization-level audits. ## Step 4 — Connect repositories Each invitee (or you, if connecting directly) imports repositories via: - **Git OAuth** — GitHub, GitLab, Azure DevOps, or Bitbucket (recommended) - **PAT or SSH** where OAuth is unavailable - **CodeDD CLI** for local-first audits — [CodeDD CLI Guide](/documentation/codedd-cli-guide) After import, CodeDD discovers files and calculates lines of code. ## Step 5 — Select audit scope and pay 1. Open **audit scope selection** for each repository. 2. Review discovered files and adjust which receive deep LLM analysis. 3. Complete checkout if scoped lines exceed your LOC budget. Purchase additional lines under **Account → Billing** before or during checkout. ## Step 6 — Start the audit 1. Invitees mark repositories **ready** when import and scope review are complete. 2. Confirm execution from the group audit workflow. 3. Track progress on the organization dashboard. When processing finishes, portfolio tabs unlock: benchmarks, technical debt, security, architecture, dependencies, development insights, and PDF export. ## Step 7 — Invite your team (optional) From **Manage Access**, invite colleagues as **Reviewer** (full report access) or **Submitter** (repository import only). Up to 10 email addresses per batch. Details: [Team Members & Access Roles](/documentation/team-members-access-roles). ## Troubleshooting | Issue | What to try | |-------|-------------| | No **Create New Audit** button | Confirm you are an **initiator** | | Import fails | Verify Git credentials — [Repository Connection & Security](/documentation/repository-connection-security) | | Checkout blocked | Top up LOC under **Account → Billing** | | Submitter cannot see KPIs | Expected — assign **Reviewer** for dashboard access | | CLI audit not appearing | Run `codedd auth status` and confirm sync completed | ## Next steps - [Common Questions — Account, Billing & Support](/documentation/common-questions-account-billing-support) - [Audit Process Overview](/documentation/audit-process-overview) - [CodeDD CLI Guide](/documentation/codedd-cli-guide) --- ## Enterprise Data API Source: https://www.codedd.ai/documentation/enterprise-data-api # Enterprise Data API The Enterprise Data API gives you programmatic, read-only access to everything CodeDD has analysed for your firm — repository audit results, findings, dependencies, architecture, and portfolio-level KPIs — so you can build your own dashboards, feed a warehouse, or attach CodeDD numbers to an investment committee pack. It is a **REST API over HTTPS returning JSON**, authenticated with an **organization-scoped API key**. Every endpoint is a `GET`; nothing in this API can change your data. **Availability:** the CodeDD **Enterprise** plan. Portfolio Monitoring-only contracts do not include it. ## Quick start ```bash # 1. Store your key outside your shell history export CODEDD_API_KEY="codedd_ent_..." # 2. Ask the API what your key can do curl -H "Authorization: Bearer $CODEDD_API_KEY" \ https://api.codedd.ai/api/v1/enterprise/ # 3. List the portfolio companies the key covers curl -H "Authorization: Bearer $CODEDD_API_KEY" \ https://api.codedd.ai/api/v1/enterprise/portcos/ ``` ## Creating a key Keys are created by an **administrator of your legal organization** in **Account → Enterprise plan → Data API keys**. 1. Click **Create API key**. 2. Give it the name of the tool that will use it — `Power BI production`, not `test` — so you know what breaks if you revoke it. 3. Choose the **scopes** it needs (see below). 4. Choose a lifetime. Every key expires; the maximum is 365 days. 5. **Copy the key immediately.** It is displayed once and cannot be retrieved again. A key belongs to the organization, not to the person who created it, so an integration keeps working when that person changes role or leaves. Up to **10 active keys** per organization. ### Scopes Grant only what the consuming tool needs. | Scope | Grants | |-------|--------| | `read:audits` | Discovery (PortCos, group audits, repository list), PortCo settings, and all single-audit data endpoints | | `read:portfolio` | All group-audit roll-ups: KPIs, technical debt, benchmark, estate map, executive summary, category breakdown, development, security findings, DORA, AI-Native, financials, repo activity, supply chain, integration assessment, KPI history | `read:audits` does **not** imply `read:portfolio`. A pipeline that only needs per-repository findings should not be able to read the firm-level roll-ups you would present to an investment committee. ### If you lose a key There is no reveal endpoint, by design: a value that can be fetched again can be fetched by the wrong person. **Rotate** the key instead. Rotation issues a replacement immediately and keeps the old key working for **24 hours**, so you can roll the secret forward on your next deploy rather than at the same instant. ## Authentication Send the key as a bearer token on every request: ``` Authorization: Bearer codedd_ent_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx ``` The key must not appear in a URL, a query string, or a client-side bundle. It reads your firm's entire audit estate; treat it like a database password and store it in your secret manager. ## Navigating your data You never have to construct an identifier. Start from the key and walk down: ``` GET /api/v1/enterprise/ → what this key can do GET /api/v1/enterprise/portcos/ → your portfolio companies GET /api/v1/enterprise/portcos/{portco_uuid}/group-audits/ → audits run for one company GET /api/v1/enterprise/group-audits/{group_audit_uuid}/audits/ → repositories in one audit GET /api/v1/enterprise/audits/{audit_uuid}/scores/ → results for one repository ``` ### PortCo endpoints No scope beyond org membership (any valid key for the firm). | Endpoint | Returns | |----------|---------| | `/api/v1/enterprise/portcos/{portco_uuid}/settings/` | Financial assumptions (`hourly_rate_usd`, `quality_threshold`, `target_coverage`, …) | | `/api/v1/enterprise/portcos/{portco_uuid}/kpi-history/` | Executive KPI values across group audits (`?group_audit_uuids=uuid1,uuid2`, optional — defaults to the 25 most recent) | ### Repository audit endpoints All require `read:audits`. | Endpoint | Returns | |----------|---------| | `/api/v1/enterprise/audits/{audit_uuid}/` | Name, status, mode, lines of code, file count, synthesis date | | `/api/v1/enterprise/audits/{audit_uuid}/scores/` | Quality scores across all assessed categories | | `/api/v1/enterprise/audits/{audit_uuid}/summaries/` | Full narrative summaries — executive, security, code quality, performance, recommendations | | `/api/v1/enterprise/audits/{audit_uuid}/flags/` | Every flag raised, with severity and location | | `/api/v1/enterprise/audits/{audit_uuid}/dependencies/` | Packages, licences, and vulnerabilities | | `/api/v1/enterprise/audits/{audit_uuid}/architecture/` | Detected technologies, components, and relationships | | `/api/v1/enterprise/audits/{audit_uuid}/development/` | Git-derived development and contribution metrics | | `/api/v1/enterprise/audits/{audit_uuid}/files/` | File statistics, extension mix, and test coverage | | `/api/v1/enterprise/audits/{audit_uuid}/executive/` | Executive dashboard payload (`?time_range=month\|quarter\|year\|all`) | | `/api/v1/enterprise/audits/{audit_uuid}/complexity/` | Cyclomatic complexity and Halstead grade distribution | | `/api/v1/enterprise/audits/{audit_uuid}/tier-benchmark/` | Tier-matched peer cohort comparison | | `/api/v1/enterprise/audits/{audit_uuid}/ai-authorship/` | Human vs AI authorship attribution | Unlike the in-app AI advisor, these return **complete payloads** — no truncated narratives, no "top 5 findings only". ### Portfolio endpoints All require `read:portfolio` and are addressed by group audit; CodeDD derives the portfolio company for you. | Endpoint | Returns | |----------|---------| | `/api/v1/enterprise/group-audits/{group_audit_uuid}/kpis/{kpi_name}/` | One KPI: `technical-debt`, `key-person`, `innovation`, `ip-security`, `scalability` | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/technical-debt/` | Full technical debt dashboard with financial impact | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/benchmark/` | Benchmark comparison | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/estate-map/` | Estate map of repositories and technologies | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/executive-summary/` | Portfolio-level narrative | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/category-breakdown/` | Subcategory score breakdown (`?category=quality`) | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/development/` | Portfolio development overview | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/security-findings/` | Issue Compass findings (paginated) | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/dora/` | DORA four-key metrics and trends | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/ai-native/` | AI-Native portfolio assessment | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/financials/` | Dollar impact modeling (`?key_person_scope=audited\|material\|root`) | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/repo-activity/` | Repository activity compass | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/supply-chain/` | Package-level license and vulnerability drilldown | | `/api/v1/enterprise/group-audits/{group_audit_uuid}/integration-assessment/` | M&A/PMI language × domain assessment | ## Coverage boundary The API exposes the **audit results and portfolio roll-ups** that reporting tools need. It does not mirror every dashboard screen — some surfaces are operational, write-only, or intentionally excluded. | Dashboard area | Available via API | Not exposed (by design) | |----------------|-------------------|-------------------------| | PortCo navigation | `/portcos/`, settings, KPI history | Member management, DORA OAuth config, remediation toggles | | Group audit discovery | Metadata, member audits | Pipeline progress, invitation workflow, scope selection | | Overview / Benchmark | KPIs, benchmark, executive summary, category breakdown, financials | Comparison matrix UI config | | Architecture | Estate map | Estate curation writes | | Development Activity | Development overview, repo activity, DORA, integration assessment | Innovation commits CSV export | | Debt & Security | Technical debt, security findings, supply chain | Saved Issue Compass selections, finding detail drawer | | AI-Native | AI-Native assessment, per-repo AI authorship | AI delivery funnel | | Single audit — Summary | Scores, summaries, executive, files (aggregates) | Per-file tree, file search | | Single audit — Files | Statistics, extension mix, test coverage, complexity | File treetable, file detail drawer | | Remediation | — | Remediation dashboard (module-gated) | | Admin / ops | — | Audit logs, comparable audits workflow | If you need a dataset that is not listed above, contact your account manager — new endpoints are added when a reporting use case is clear and the underlying data is stable. ## Response format Every successful response uses the same envelope: ```json { "status": "success", "data": { "portcos": [ { "portco_uuid": "...", "name": "Acme GmbH" } ] }, "meta": { "api_version": "1", "correlation_id": "3f2b...", "next_cursor": null, "total_count": 12 } } ``` Quote the `correlation_id` when contacting support — it identifies the exact request in our logs. ### Pagination List endpoints are paginated. Pass `limit` (default 50, maximum 100) and follow `meta.next_cursor` until it is `null`: ```python import requests session = requests.Session() session.headers['Authorization'] = f'Bearer {API_KEY}' BASE = 'https://api.codedd.ai/api/v1/enterprise' def fetch_all(path, key): """Every page of a list endpoint.""" items, cursor = [], None while True: params = {'limit': 100} if cursor: params['cursor'] = cursor payload = session.get(f'{BASE}{path}', params=params, timeout=30).json() items.extend(payload['data'][key]) cursor = payload['meta']['next_cursor'] if not cursor: return items portcos = fetch_all('/portcos/', 'portcos') ``` Treat the cursor as opaque. It is not an offset, and building your own would break the moment ordering changes. ### Caching with ETags Audit results are written once and then do not change, so repeat reads should be conditional. Repository endpoints return an `ETag`; send it back as `If-None-Match` and a `304 Not Modified` costs you nothing: ```bash curl -H "Authorization: Bearer $CODEDD_API_KEY" \ -H 'If-None-Match: "a1b2c3d4..."' \ https://api.codedd.ai/api/v1/enterprise/audits/$AUDIT/scores/ ``` An audit still being analysed has no `ETag`, because its data is still moving. ## Errors Failures use a stable `error_code` you can branch on. Messages may be reworded; codes will not change within v1. | HTTP | `error_code` | Meaning and what to do | |------|--------------|------------------------| | 400 | `invalid_parameter` | A parameter failed validation. Fix the request. | | 401 | `invalid_api_key` | Missing, malformed, or unknown key. Check the header. | | 401 | `api_key_expired` | Past its expiry. Rotate it. | | 401 | `api_key_revoked` | Revoked by an administrator. Ask for a new one. | | 403 | `insufficient_scope` | The key lacks the scope this endpoint needs. Mint a key with it. | | 403 | `enterprise_plan_required` | No active Enterprise plan on the organization. | | 404 | `resource_not_found` | The resource does not exist **or** is not covered by your key. These are deliberately indistinguishable. | | 409 | `audit_in_progress` | The audit has not finished, so this result set does not exist yet. Retry later. | | 429 | `rate_limit_exceeded` | Honour the `Retry-After` header. | | 500 | `internal_error` | Our fault. Retry with backoff; quote the `correlation_id`. | ```json { "status": "error", "error_code": "insufficient_scope", "message": "This API key does not have the \"read:portfolio\" scope.", "correlation_id": "9c1f..." } ``` ### Rate limits | Bucket | Limit | |--------|-------| | Standard endpoints, per key | 120 requests/minute | | Heavy endpoints (findings, dependencies, files, all portfolio aggregates), per key | 30 requests/minute | | All endpoints, per organization | 600 requests/minute | Limits are per **key**, not per IP address, so one pipeline cannot throttle another that happens to share an egress address. On a `429`, wait for `Retry-After` seconds; do not retry immediately. ## Building a reporting pipeline A complete nightly extract: ```python import os import requests API_KEY = os.environ['CODEDD_API_KEY'] BASE = 'https://api.codedd.ai/api/v1/enterprise' session = requests.Session() session.headers['Authorization'] = f'Bearer {API_KEY}' def get(path, **params): response = session.get(f'{BASE}{path}', params=params, timeout=60) if response.status_code == 429: raise RuntimeError(f"Rate limited; retry after {response.headers['Retry-After']}s") if response.status_code == 409: return None # Audit still running; nothing to extract yet. response.raise_for_status() return response.json()['data'] rows = [] for portco in get('/portcos/', limit=100)['portcos']: group_audits = get(f"/portcos/{portco['portco_uuid']}/group-audits/", limit=100) for group_audit in group_audits['group_audits']: group_uuid = group_audit['group_audit_uuid'] debt = get(f'/group-audits/{group_uuid}/kpis/technical-debt/') if debt is None: continue for audit in get(f'/group-audits/{group_uuid}/audits/', limit=100)['audits']: # Skip repositories whose analysis has not produced results. if not audit['is_completed']: continue scores = get(f"/audits/{audit['audit_uuid']}/scores/") rows.append({ 'portco': portco['name'], 'group_audit': group_audit['name'], 'repository': audit['name'], 'lines_of_code': audit['lines_of_code'], 'scores': scores['scores'], 'portfolio_technical_debt': debt['data'], }) print(f'Extracted {len(rows)} repositories') ``` ### Recommended practice - **Store the key in a secret manager**, never in source control, a notebook, or a dashboard definition. - **Rotate on a schedule** — quarterly is a reasonable default — and always after someone with access leaves. - **Use one key per consuming system.** When something misbehaves you can revoke it without taking down every other integration. - **Grant the narrowest scope** the tool needs. - **Cache with `ETag`s** and only re-read what changed. A completed audit's results never change. - **Retry with exponential backoff** on `429` and `5xx`; never retry a `4xx` other than `429`. - **Check `is_completed`** before extracting results, and treat `409 audit_in_progress` as "come back later", not as an error. ## Security and auditability - Keys are stored only as SHA-256 digests. Neither a database dump nor a query log yields a usable credential. - Every request is recorded against the key — method, route, resource, status, and outcome — including refusals. Ask support for an access export if you need to evidence who read what. - Access is re-checked on every request against your live contract, the key's scopes, and the resource's ownership. Nothing is inherited from when the key was created. - Responses are marked `Cache-Control: private, no-store` so audit findings do not linger in an intermediary cache. - All traffic must be HTTPS. ## Frequently asked **Can I write data through this API?** No. Every endpoint is read-only. Audits are started in the app or through the [CodeDD CLI](/documentation/codedd-cli). **Can I use it with an MCP client?** The API is the foundation for MCP access to CodeDD data. Contact your account manager about current MCP availability. **What happens when our contract lapses?** The API closes within a minute, and existing keys stop working. They resume if the contract is reinstated and has not expired. **Can a key read another firm's data?** No. A key is bound to one legal organization and can only reach portfolio companies, audits, and results that belong to it. Anything else returns `404`. **Do invited team members access the API with their login?** No. The data endpoints do not accept session JWTs — only an organization API key (`Authorization: Bearer codedd_ent_...`). Invited members do not get automatic API access; an **organization administrator** must create a key and share it with the tool (or person) that needs it. Keys are org-scoped: whoever holds the key can read the firm's entire entitled audit estate, not just the subset one member sees in the web UI. **Who can create or revoke keys?** Only administrators of the **legal organization** (the Enterprise billing entity), via Account → Enterprise plan → Data API keys, while on an active full Enterprise contract. --- ## File Discovery & Indexing Source: https://www.codedd.ai/documentation/file-discovery-indexing # File Discovery & Indexing After your repository is connected, CodeDD scans the full tree, calculates metrics, and prepares files for analysis. Everything in the audit cache is encrypted at rest. ## What gets scanned - Source code (100+ languages) - Configuration (Docker, Kubernetes, CI/CD, env files) - Infrastructure-as-Code (Terraform, CloudFormation) - Documentation and database schemas ## What gets excluded - Binary files and build outputs (`dist`, `build`, `target`) - Version control (`.git`) - Dependencies (`node_modules`, `vendor`, `venv`) - Symlinked directories (to avoid duplicate counts) ## File categorization Files are automatically typed as source code, configuration, infrastructure, security-related, or documentation. This drives scope defaults and architecture detection. ## Metrics collected For each file: - **Lines of code** — non-empty code lines, excluding comments and whitespace - **Documentation lines** — comments, docstrings, README content - **Git metadata** — last modified, contributors, commit frequency, churn Folder-level aggregates roll up LOC, file counts, and language breakdowns for dashboard views. ## Audit scope After discovery, files are marked for deep analysis. You can review and adjust scope in the UI before the audit runs — include or exclude paths, or accept auto-selection defaults that prioritize source and security-relevant files. Not every indexed file receives LLM analysis. Scope focuses effort on files that matter for due diligence. ## Encryption As files are indexed, content is encrypted in the audit cache using Fernet (AES-128-CBC + HMAC). Plaintext exists only transiently in memory during later analysis stages. → [Data Encryption at Rest](/documentation/data-encryption-at-rest) ## Progress tracking During import, the UI shows files processed, percentage complete, and processing speed so you can monitor large repositories. ## Error handling Individual file failures (permissions, encoding) are logged and skipped — they do not stop the audit. Large files are handled via streaming. ## What happens next Scoped, encrypted files proceed to LLM file analysis, dependency scanning, architecture mapping, and consolidation. ## Next steps - [AI-Powered File Analysis](/documentation/ai-powered-file-analysis) - [Data Encryption at Rest](/documentation/data-encryption-at-rest) - [Secure Data Deletion](/documentation/secure-data-deletion) --- ## Compliance & Certifications Source: https://www.codedd.ai/documentation/compliance-certifications # Compliance & Certifications CodeDD maintains a security assurance program aligned with **ISO/IEC 27001** and **SOC 2 Type II**. ## Framework alignment | Framework | Status | |-----------|--------| | ISO/IEC 27001 | Certified | | SOC 2 Type II | Certified | | GDPR | Privacy Policy, DPA, SCCs available | | NIST-aligned practices | Encryption, access control, logging, incident response | Independent penetration tests are conducted regularly. Compliance artifacts are available under NDA on request. ## Data privacy **GDPR** — lawful basis in Privacy Policy and DPA; data minimization (source code deleted after processing); right to erasure; Standard Contractual Clauses for international transfers; sub-processor list in the DPA. Contact compliance@codedd.ai for jurisdiction-specific questions (CCPA, UK GDPR, etc.). ## Technical controls - **Encryption** — Fernet at rest, TLS in transit, key rotation support → [Data Encryption at Rest](/documentation/data-encryption-at-rest) - **Access control** — role-based access, two-factor authentication, CLI tokens with expiry and revocation - **Secure deletion** — 3-pass overwrite after audit → [Secure Data Deletion](/documentation/secure-data-deletion) - **Monitoring** — application and infrastructure logging with incident response procedures ## Data residency Primary hosting: **IONOS data centers in Germany (EU)**. Stated in the Privacy Policy, Terms, DPA, and Security page. Enterprise contracts may specify alternate arrangements. ## Sub-processors AI inference providers process source code transiently under commercial agreements with data-protection terms. CodeDD does not use customer code to train AI models. Sub-processors are listed in the DPA. ## Shared responsibility **CodeDD provides:** secure processing, encryption and deletion controls, platform access management, compliance documentation on request. **Customers provide:** read-only repository access with appropriate scoping, user access management within their organization, notification of special regulatory requirements. ## Available documentation **On request (NDA may apply):** SOC 2 report, ISO certificate (when available), penetration test summaries, DPA, security policies, sub-processor list. **Public:** [Privacy Policy](https://www.codedd.ai/privacy) · [Terms of Service](https://www.codedd.ai/terms) · [Security page](https://www.codedd.ai/security) Contact: **compliance@codedd.ai** ## Next steps - [Data Encryption at Rest](/documentation/data-encryption-at-rest) - [Secure Data Deletion](/documentation/secure-data-deletion) - [Repository Connection & Security](/documentation/repository-connection-security) --- ## Team Members & Access Roles Source: https://www.codedd.ai/documentation/team-members-access-roles # Team Members & Access Roles Portfolio organizations support shared access for investment teams, advisors, and portfolio company engineers. Each member gets one of three roles. ## Organization roles | Role in the UI | Who gets it | What they can do | |----------------|-------------|------------------| | **Initiator** | Organization creator (owner) | Full access to dashboards, audits, and settings. Only initiators can invite or remove members and create organization audits. | | **Reviewer** | Colleagues invited as Reviewer | Full access to organization audits, repositories, and portfolio reports. Cannot invite or remove members. | | **Submitter** | Colleagues invited as Submitter | Can connect repositories and submit them for audit. No access to portfolio dashboard analytics or executive summaries. | CodeDD administrators retain full access across organizations for support. ## Inviting members Only an **initiator** can invite people. 1. Open your organization dashboard. 2. Go to **Invite Members** (Manage Access). 3. Choose **Reviewer** or **Submitter**. 4. Enter up to **10 email addresses** per batch and send. **Existing users** receive a link to the organization dashboard. **New users** get a registration link; their role applies after signup. Pending invitations show as **pending** until registration completes. ## Removing members Only initiators can remove members. Initiators cannot be removed through the UI. Removed users lose organization access immediately. ## Role comparison | Capability | Initiator | Reviewer | Submitter | |------------|-----------|----------|-----------| | View portfolio dashboards & reports | Yes | Yes | No | | View individual audit findings | Yes | Yes | No | | Submit / import repositories | Yes | Yes | Yes | | Invite or remove members | Yes | No | No | | Create organization audits | Yes | No | No | ## Audit-level roles Organization roles are separate from **audit invitation** roles (inviter, invitee, auditor) used when a specific audit is shared directly. For group audits, initiators act on the auditor side; reviewers and submitters typically participate on the invitee side for repository submission. ## Need help? Use **Reviewer** for colleagues who need to read reports. Use **Submitter** for external engineers who should only upload code. Contact [CodeDD support](mailto:info@codedd.ai) to transfer initiator ownership. ## Related guides - [Setting Up a Portfolio Organization](/documentation/setting-up-a-portfolio-organization) - [Common Questions — Account, Billing & Support](/documentation/common-questions-account-billing-support) --- ## AI-Powered File Analysis Source: https://www.codedd.ai/documentation/ai-powered-file-analysis # AI-Powered File Analysis CodeDD's file analysis stage reviews each scoped file for quality, security, and maintainability — surfacing issues that pattern-only scanners often miss. ## Beyond traditional SAST Traditional static analysis relies on pattern matching and syntax rules. CodeDD adds **semantic understanding**: business logic review, intent-based risk detection, and context from file type and role in the system. Security findings go through a **validation step** that checks for supporting evidence before they affect your score. Optional SonarQube integration adds language-specific static rules when enabled. ## What gets analyzed **Application code** — business logic, APIs, database access, auth, validation, error handling. **Infrastructure** — Dockerfiles, Kubernetes manifests, CI/CD configs, IaC templates. **Configuration** — app settings, environment files, third-party integrations. ## File selection Deep analysis runs on files marked in **audit scope**. Priority goes to security-critical paths, core business logic, recently modified files, and high-complexity code. Excluded: auto-generated code, minified files, binaries, and test fixtures. ## Per-file assessment Each analyzed file receives structured scores across dimensions you see in the dashboard: | Area | What it covers | |------|----------------| | Code quality | Readability, consistency, modularity, maintainability, technical debt | | Functionality | Completeness, edge cases, error handling | | Performance | Efficiency, scalability, resource use | | Security | Input validation, data handling, authentication | | Standards | Best practices, design patterns, complexity | Findings include severity (Green / Yellow / Orange / Red), confidence level, and specific remediation guidance. ## Security findings Common categories: injection flaws, XSS vectors, auth bypasses, insecure cryptography, exposed secrets, sensitive data handling. Each security flag carries a confidence score from the validation pipeline. Findings at or above the **80% confidence threshold** are treated as validated and actionable; lower-confidence items are marked inconclusive for manual review in the Security / Flags tab. ## Dependency analysis CodeDD scans package manifests and imports across major ecosystems (npm, pip, Maven, Go modules, Cargo, Composer, .NET, and others). For each dependency: - Version and known CVEs (NVD, GitHub Advisory, OSV) - Severity and patch availability - License type and compliance flags Portfolio views include a **Supply Chain Vulnerabilities** panel and **License Compliance** section. ## Complexity metrics Cyclomatic complexity is calculated per function using Radon/Lizard. High-complexity functions (typically 20+) are flagged as refactoring candidates and contribute to technical debt signals. ## SonarQube (optional) When enabled, SonarScanner runs in an isolated container against a temporary workspace. Results are merged with AI findings. The workspace is deleted after analysis. ## Data privacy After analysis: - **Stored:** file paths, metrics, findings, dependency lists, complexity scores - **Never stored:** source code content, detected secrets (flagged but not persisted), PII from comments ## Next steps - [Cross-File Contextualization](/documentation/cross-file-contextualization) - [Audit Consolidation & Risk Scoring](/documentation/audit-consolidation-risk-scoring) - [Architecture Analysis & Mapping](/documentation/architecture-analysis-mapping) --- ## Cross-File Contextualization Source: https://www.codedd.ai/documentation/cross-file-contextualization # Cross-File Contextualization Individual file analysis catches local issues. Contextualization connects them — mapping how components interact, where patterns break down, and where systemic risk concentrates. This runs after [Architecture Analysis & Mapping](/documentation/architecture-analysis-mapping) and before consolidation. ## Why it matters File-by-file scanners miss cross-cutting problems: inconsistent auth checks, data flowing without sanitization, duplicated logic across modules, and architectural anti-patterns that only appear at system level. ## Domain mapping CodeDD groups files into logical domains based on directory structure, naming, imports, and framework patterns: - **Frontend** — UI components, routing, client state - **Backend** — APIs, business logic, middleware - **Database** — schemas, migrations, ORM models - **Infrastructure** — containers, orchestration, CI/CD, config - **Testing** — unit, integration, and E2E tests Per domain, you see LOC, file count, complexity, test coverage, and issue density. ## What contextualization checks **Cross-domain flows** — how frontend, backend, and data layers connect; whether boundaries are respected. **Security consistency** — auth mechanisms, protected vs. unprotected endpoints, credential handling across files. **Data flow** — where sensitive data originates, how it moves, whether validation and sanitization happen at each boundary. **Systemic gaps** — missing input validation, inadequate error handling, absent rate limiting, insufficient test coverage on critical paths. **Knowledge concentration** — contributor distribution per domain, bus-factor signals, abandoned or high-churn areas. ## Portfolio-level context For group audits, contextualization also surfaces cross-repository patterns: shared dependencies, common vulnerability types, and technology stack consistency across the portfolio. ## What gets stored Domain names, file-to-domain mappings, metric summaries, architectural patterns, gap findings, and recommendations. Source code and snippets are not stored. ## Next steps - [Audit Consolidation & Risk Scoring](/documentation/audit-consolidation-risk-scoring) - [Recommendations Generation](/documentation/recommendations-generation) - [Architecture Analysis & Mapping](/documentation/architecture-analysis-mapping) --- ## Architecture Analysis & Mapping Source: https://www.codedd.ai/documentation/architecture-analysis-mapping # Architecture Analysis & Mapping Most codebases lack up-to-date architecture documentation. CodeDD reverse-engineers structure from the code itself — producing an interactive map of components, technologies, relationships, and data flows. ## What you get **Interactive architecture diagram** — color-coded nodes (frontend, backend, database, infrastructure) with directional edges showing data flow. Click a component to see associated files. Zoom, pan, and filter in the dashboard. **Technology inventory** — languages, frameworks, databases, external services, and infrastructure tools detected across the repository, with version information where available. **Architectural assessment** — detected patterns (monolith, layered, microservices-ready), scalability indicators, coupling analysis, and risk highlights (single points of failure, outdated stack, missing abstractions). **Test coverage by component** — which parts of the system are well-tested and which are not. ## How it works (at a high level) CodeDD combines three approaches: 1. **Pattern detection** — scans manifests, configs, and source structure to identify technologies, dependency files, deployment configs, and database schemas across major programming languages and frameworks. 2. **AI classification** — analyzes key components to determine role, tech stack, and architectural implications. 3. **Graph synthesis** — builds the interactive diagram with components organized into layers (code, data/communication, deployment) and mapped relationships. Confidence scores accompany findings so you can assess reliability — standard stacks score higher; heavily customized systems may have lower confidence on inferred relationships. ## Typical domains detected - Application services (APIs, business logic, workers) - Frontend clients (React, Vue, Angular, mobile) - Data stores (PostgreSQL, MySQL, MongoDB, Redis, message queues) - Infrastructure (Docker, Kubernetes, CI/CD, reverse proxies) - External integrations (payment, email, storage, monitoring) ## Use cases **Investors** — Is the architecture modern? Are there scaling bottlenecks or single points of failure? What would integration or replatforming cost? **CTOs** — Onboarding map for new engineers. Data-driven refactoring priorities. Technology audit inventory. **M&A advisors** — Stack compatibility with acquirer tech, knowledge transfer complexity, replatforming scope. ## Limitations Architecture analysis is **static** — it reads code, not runtime behavior. - Data flows are inferred, not measured live - Dynamically loaded components may be missed - Microservices split across multiple repos need separate audits per repo - Custom internal frameworks may not be fully recognized Security vulnerability scanning is a separate feature covered in [AI-Powered File Analysis](/documentation/ai-powered-file-analysis). ## Next steps - [Cross-File Contextualization](/documentation/cross-file-contextualization) - [Audit Consolidation & Risk Scoring](/documentation/audit-consolidation-risk-scoring) - [AI-Powered File Analysis](/documentation/ai-powered-file-analysis) --- ## Audit Consolidation & Risk Scoring Source: https://www.codedd.ai/documentation/audit-consolidation-risk-scoring # Audit Consolidation & Risk Scoring Consolidation turns file-level findings, dependency data, and architecture insights into portfolio-ready scores and reports. ## Code Health Score The **Code Health Score (0–100)** is the primary metric for investment decisions. It reflects composite debt across three weighted components: ``` Composite Debt = (Code Quality Debt × 50%) + (Validated Security Debt × 30%) + (Test Coverage Debt × 20%) Code Health Score = max(0, 100 − Composite Debt) ``` | Component | Weight | What it measures | |---|---|---| | Code Quality | 50% | Maintainability, readability, modularity, complexity | | Validated Security | 30% | Confirmed security exposure from validated findings and CVEs | | Test Coverage | 20% | Gap between actual coverage and 100% | Your organization's **quality threshold** and **target coverage** (typically 80%) affect remediation cost estimates in the dashboard — not the health score itself. ## Component detail ### Code quality Derived from the overall quality score computed during file analysis: readability, modularity, maintainability, redundancy, and technical debt. For older audits without an overall quality score, technical debt score is used as fallback. ``` Code Quality Debt = 100 − Code Quality Score ``` ### Validated security Based on validated security findings and dependency CVEs, normalized per repository so a large portfolio is not automatically penalized more than a single repo. Severity penalties use graduated weights — critical issues have outsized impact compared to medium findings. Eliminating one critical vulnerability per repo moves the score more than clearing several medium issues. ### Test coverage ``` Test Coverage Debt = max(0, 100 − average coverage %) ``` Averaged across repositories with coverage data. ### Supplementary metrics Shown in dashboards but **not** in the Code Health Score formula: - **Maintenance indicators** — validated Orange/Red flag density per 1,000 LOC - **Supply chain score** — dependency vulnerability exposure - **Per-repository security breakdown** — critical/high/medium counts ## Score bands | Band | Score | Typical profile | |------|-------|-----------------| | **Good** | 67–100 | Low debt, minimal critical/high vulns, coverage near target | | **Fair** | 33–66 | Moderate debt in one or more areas; manageable with a plan | | **Poor** | 0–32 | High debt, critical vulns present, low coverage | ## Portfolio aggregation For group audits: - **Code quality** — LOC-weighted average across repositories - **Security** — severity counts summed, then averaged per repo before scoring - **Test coverage** — arithmetic average across repos with data - **Maintenance** — portfolio-wide flag density per 1,000 LOC Compare repositories within a portfolio to identify outliers and leaders. ## What you see in the dashboard - Code Health Score and KPI strip with component drill-down - Executive summary with top findings and strengths - Security flags, supply-chain panel, architecture map - Issue Compass for prioritized remediation - Compare previous audits (single-repo and group) - PDF export for IC materials JIRA / GitHub issue creation from findings is **planned**. ## Re-audit and trends Re-run audits and use compare-audits to track score changes, resolved issues, and coverage improvements over time. ## Next steps - [Recommendations Generation](/documentation/recommendations-generation) - [Data Encryption at Rest](/documentation/data-encryption-at-rest) - [Secure Data Deletion](/documentation/secure-data-deletion) --- ## Recommendations Generation Source: https://www.codedd.ai/documentation/recommendations-generation # Recommendations Generation Recommendations are produced during audit consolidation — turning validated findings into prioritized, actionable guidance tied to your codebase. ## What recommendations cover **Security remediation** — specific vulnerabilities with severity, affected location, and fix guidance (e.g. parameterized queries instead of string concatenation, adding auth checks to exposed endpoints). **Dependency updates** — outdated packages with known CVEs, target versions, and breaking-change notes where relevant. **Technical debt** — high-complexity functions to refactor, duplication to consolidate, missing error handling. **Architecture** — coupling issues, missing abstraction layers, modularity improvements. **Testing** — coverage gaps on critical paths, suggested test types (unit, integration, security). Each recommendation includes priority tier, estimated effort, and rationale. ## Prioritization Recommendations are ranked by severity, confidence, and business impact: - **P0 — Critical:** active security vulnerabilities, data breach risks, production-breaking issues - **P1 — High:** high-severity vulns, important dependency updates, major architecture issues - **P2 — Medium:** code quality improvements, moderate refactoring - **P3 — Low:** style consistency, documentation, minor optimizations Related fixes are grouped where dependencies exist (e.g. framework upgrade before API migration). ## Where to use them **Web dashboard** — prioritized list with drill-down to findings and affected files. **Issue Compass** — focused remediation slices you can save and work through systematically. **PDF export** — executive summary with top recommendations. **CLI** — `codedd fix` for terminal-guided resolution after audit completion. **Exports** — JSON/CSV for offline planning; API access for programmatic retrieval. ## Tracking progress Re-run audits and use **compare-audits** to measure: - Issues resolved vs. new findings - Code Health Score change - Test coverage delta - Dependency updates completed Audits are on your cadence — not automated continuous monitoring. ## Planned - JIRA / GitHub issue auto-creation (UI currently shows "coming soon") - CI/CD webhook triggers for re-audit ## Next steps - [Audit Consolidation & Risk Scoring](/documentation/audit-consolidation-risk-scoring) - [CodeDD CLI Guide](/documentation/codedd-cli-guide) - [Secure Data Deletion](/documentation/secure-data-deletion) ---