Compliance dossier — training data (AI Act, Article 10)
Automatically generated skeleton from the dataset analysis. The "to be completed by the expert" blocks require human review (provenance, compliance judgment, mitigation measures).
1. Identity & purpose
- System / model: _____
- Dossier version: _____
- Analysis date: _____
- Dataset:
credit - Intended purpose: _____
_To be completed by the expert — intended purpose and what the data is meant to represent_
2. Provenance & legal basis
Origin of each source, licences/contracts, and — for personal data — initial purpose and GDPR legal basis.
⚠️ Potential personal data detected: email, full_name — a GDPR legal basis is required for these columns.
_To be completed by the expert — provenance per source (separate flow §16), licences, GDPR legal basis_
3. Composition _(filled automatically)_
- Volume: 510 rows, 10 columns.
- Variables:
| Column | Type |
id | BIGINT |
age | BIGINT |
gender | VARCHAR |
postal_code | BIGINT |
income | BIGINT |
employment_years | DOUBLE |
loan_amount | BIGINT |
email | VARCHAR |
full_name | VARCHAR |
defaulted | BIGINT |
- Declared sensitive attributes:
gender,postal_code,age.
- Detected PII columns:
email,full_name.
_To be completed by the expert — geographic / contextual / behavioural scope_
4. Preparation _(automatic observations)_
- Duplicates: 10 duplicate row(s) → deduplication step to be traced.
- Missing values:
income(25) → handling strategy to be documented.
_To be completed by the expert — transformation log: collection, cleaning, labelling, enrichment, aggregation_
5. Quality _(filled automatically)_
- Completeness: 25 missing value(s) (
income). - Accuracy / outliers:
income(3). - Uniqueness: 10 duplicate(s).
_To be completed by the expert — representativeness and relevance to the purpose_
6. Bias
Automatic analysis
| Characteristic | Measure | Target | Verdict |
| Completeness | 99.5 % | ≥ 95.0 % | compliant |
| Uniqueness | 98.0 % | ≥ 99.0 % | to address |
| Validity | 99.9 % | ≥ 98.0 % | compliant |
gender — « defaulted »: disparate impact (ratio) 2.14 · statistical parity (gap) 16.2 % ⚠️
| Group | Count | « defaulted » rate |
F | 250 | 30.4 % |
M | 260 | 14.2 % |
postal_code — « defaulted »: disparate impact (ratio) 1.83 · statistical parity (gap) 14.0 % ⚠️
| Group | Count | « defaulted » rate |
93200 | 100 | 31.0 % |
59000 | 107 | 24.3 % |
13008 | 103 | 19.4 % |
69003 | 88 | 19.3 % |
75001 | 112 | 17.0 % |
Judgment & measures
_To be completed by the expert — impact on health/safety/fundamental rights and detection/prevention/mitigation measures_
7. Gaps & limitations
Detected points to examine as potential gaps:
- [medium] duplicates
- [low] missing —
income - [low] outliers —
income - [high] pii —
email - [high] pii —
full_name - [high] bias —
gender - [medium] bias —
postal_code
_To be completed by the expert — identified and addressed gaps; out-of-scope uses_
8. Governance & traceability
- Analysis tool version:
0.4.1
- Dataset fingerprint (schema + volume):
d03427e41c5697a1
_To be completed by the expert — responsibilities, audit log, versioning, maintenance/updates_