syncova-backup/SYNCOVA_IMPLEMENTATION_PLAN.md
Jerrit Fritzsche 610719c316
Some checks failed
CI / Backend (Go) (push) Failing after 3m7s
CI / Frontend (React/TypeScript) (push) Successful in 37s
CI / Sicherheitsprüfungen (push) Successful in 44s
Syncova Backups V1
Enterprise-Backup-, Recovery-, Verification-, Security- und
Monitoring-Plattform fuer Proxmox VE, Windows, Linux und Dateisysteme.

Der Leitsatz, der fast jede Entscheidung erklaert: Ein Backup gilt erst als
vertrauenswuerdig, wenn Integritaet geprueft und Wiederherstellbarkeit
nachgewiesen wurde. Deshalb steigt ein Wiederherstellungspunkt erst nach einem
tatsaechlich durchgefuehrten Restore-Test auf "recoverable", und Unbekanntes
geht in keine Bewertung als "gut" ein.

Umfang (Phasen 0-23):

- Repository Engine: inhaltsadressierte Bloecke, atomares Commit-Protokoll,
  Katalogaufbau allein aus den Manifesten — ohne Datenbank
- Backup Engine: inhaltsabhaengiges Chunking, Deduplizierung trotz
  Verschluesselung, zstd, AES-256-GCM, Streaming mit Gegendruck
- Agenten fuer Windows und Linux mit Auftragsabholung (Pull-Modell)
- Proxmox-Provider mit beiden Zugriffswegen auf die Sicherungsarchive
- Scheduler, Recovery Engine mit Pruefpunkt, Verification, Unveraenderlichkeit
- Weboberflaeche, Kennzahlen, Meldungen, Berichte, Security Center,
  Ransomware-Heuristik (meldet, handelt nie)
- Disaster Recovery, Haertung, Leistungsmessung, Chaos Testing
- Eingefrorene Vertraege fuer API, Migrationen, Backup-Format und Repository
- Auslieferungspaket fuer linux/amd64, linux/arm64 und windows/amd64

Nicht enthalten und als solches gekennzeichnet: Kapazitaetsprognose, Backup
Copy, Changed Block Tracking bei Proxmox, erweiterte Attribute und ACLs.

Gebaut, aber nie auf echter Hardware gefahren: der Windows-Dienst, die
systemd-Einheit und der verpflichtende Proxmox-Meilenstein — ob eine
wiederhergestellte VM startet, ist ungeprueft. Einzelheiten in CHANGELOG.md
und docs/release-candidate.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 09:10:54 +02:00

836 lines
12 KiB
Markdown

# Syncova Backups V1 — Implementation Plan
## 1. Objective
Build Syncova V1 as a production-oriented backup and recovery platform with a staged implementation process.
Do not attempt to implement all functionality simultaneously.
Each phase ends with:
- working software
- automated tests
- documentation
- security review
- measurable acceptance criteria
## 2. Phase 0 — Product Foundation
### Deliverables
- repository structure
- coding standards
- architecture documentation
- CI pipeline
- dependency management
- local development environment
- PostgreSQL
- migration framework
- basic frontend shell
- backend health endpoint
### Suggested repository
```text
syncova/
apps/
api/
web/
agent/
worker/
packages/
backup/
repository/
crypto/
providers/
metrics/
auth/
migrations/
docs/
tests/
deployment/
```
### Exit Criteria
- Backend starts.
- Frontend starts.
- Database connects.
- Migration runs.
- CI executes tests.
- No hardcoded credentials.
---
# 3. Phase 1 — Identity and Security Foundation
Implement:
- users
- roles
- permissions
- login
- sessions/tokens
- password hashing
- TOTP MFA
- audit logging
- TLS configuration
- secret abstraction
### Exit Criteria
- Admin can log in.
- MFA works.
- Unauthorized users cannot access protected APIs.
- Privileged operations are audited.
- Passwords are never stored plaintext.
---
# 4. Phase 2 — Repository Engine
This is the first major engineering milestone.
Implement:
- repository abstraction
- local filesystem repository
- hardened repository mode
- chunk storage
- manifests
- atomic commit
- repository catalog
- integrity metadata
- repository health
- repository scan
### Tests
- write chunk
- read chunk
- detect corruption
- interrupted write
- restart
- catalog rebuild
### Exit Criteria
A repository can be created, written, scanned, restarted, and rebuilt without PostgreSQL containing the only copy of backup metadata.
---
# 5. Phase 3 — Backup Format
Implement versioned Syncova format.
Must contain:
- header
- source metadata
- backup metadata
- manifest
- chunk references
- integrity information
- encryption metadata
- completion marker
### Tests
- serialize
- deserialize
- version compatibility
- invalid manifest
- missing chunk
- corrupted chunk
- incomplete backup
---
# 6. Phase 4 — Backup Engine Core
Implement streaming pipeline:
```text
Read
-> Chunk
-> Hash
-> Dedup
-> Compress
-> Encrypt
-> Write
-> Manifest
-> Commit
```
Implement:
- bounded worker pools
- backpressure
- checkpoints
- cancellation
- retry
- progress reporting
### Metrics
Collect:
- bytes processed
- bytes written
- throughput
- compression ratio
- dedup ratio
- duration
### Exit Criteria
A large test dataset can be backed up without loading the dataset into memory.
---
# 7. Phase 5 — Windows Agent
Implement:
- Windows service
- enrollment
- heartbeat
- file discovery
- file backup
- restore
- secure communication
- local logging
### Tests
- service restart
- network interruption
- permission errors
- large files
- many small files
- incremental backup
### Exit Criteria
A Windows system can be backed up and restored reliably.
---
# 8. Phase 6 — Linux Agent
Implement:
- systemd service
- enrollment
- heartbeat
- filesystem backup
- restore
- secure communication
Test:
- permissions
- symlinks
- special files as applicable
- interrupted jobs
- incremental operation
---
# 9. Phase 7 — Proxmox Provider
Implement provider interface and Proxmox implementation.
### Discovery
- cluster
- hosts
- VMs
- VM storage
- configuration
### Backup
- VM metadata
- VM disks
- snapshots where appropriate
- changed blocks where available
- consistency state
### Restore
- original host
- alternate host
- new VM
### Exit Criteria
A test Proxmox environment can:
```text
Discover VM
-> Backup VM
-> Verify
-> Delete test VM
-> Restore VM
-> Boot VM
-> Validate VM
```
This is a mandatory end-to-end milestone.
---
# 10. Phase 8 — Scheduler and Job Management
Implement:
- jobs
- schedules
- priorities
- concurrency
- retries
- backoff
- job dependencies
- bandwidth limits
- maintenance windows
### UI
Create the backup wizard:
1. Name
2. Source
3. Schedule
4. Repository
5. Retention
6. Security
7. Verification
8. Notifications
9. Review
10. Create
---
# 11. Phase 9 — Recovery Engine
Implement:
- file restore
- folder restore
- disk restore
- full system restore where supported
- Proxmox VM restore
- restore validation
- restore sessions
- checkpoints
- progress
### Safety
Restores to production must require explicit confirmation.
---
# 12. Phase 10 — Verification and Recovery Assurance
Implement:
- chunk verification
- manifest verification
- chain validation
- repository integrity scans
- restore-point validation
- automated recovery tests
- Recovery Assurance score
### Score Inputs
- backup freshness
- verification
- restore test
- RPO
- RTO
- immutability
- offsite
- encryption
- repository health
- anomalies
### Exit Criteria
A backup can be objectively classified as:
- successful
- verified
- recoverable
- failed
- corrupted
---
# 13. Phase 11 — Immutability
Implement hardened repository protections.
Requirements:
- retention lock
- immutable-until metadata
- restricted delete
- audit
- cleanup protection
Where the storage layer supports stronger WORM/object-lock behavior, integrate it.
Do not claim true immutability if the underlying storage cannot enforce it.
---
# 14. Phase 12 — Web Dashboard
Build the main UI.
### Pages
- Dashboard
- Backup Jobs
- Protected Systems
- Recovery
- Recovery Points
- Repositories
- Proxmox
- Agents
- Alerts
- Events
- Security Center
- Reports
- Users
- Roles
- Settings
### Dashboard widgets
- protected systems
- backup success rate
- failed jobs
- critical alerts
- storage
- capacity forecast
- Recovery Assurance
- Security Score
- RPO compliance
- repository health
---
# 15. Phase 13 — Metrics and Graphs
Implement metrics storage and chart APIs.
Required graphs:
- backup success/failure
- duration
- throughput
- processed/written data
- storage growth
- deduplication
- compression
- repository capacity
- agent resources
- RPO
- recovery duration
- verification results
Time ranges:
- 1h
- 24h
- 7d
- 30d
- 90d
- 1y
- custom
---
# 16. Phase 14 — Alerts and Notifications
Implement:
- alert rules
- severity
- acknowledge
- resolve
- email
- webhook
Initial rules:
- backup failure
- repeated backup failure
- repository unavailable
- repository >80%
- repository >90%
- agent offline
- RPO violation
- verification failure
- certificate expiry
- possible ransomware activity
- immutable copy missing
---
# 17. Phase 15 — Security Center
Implement:
- Security Score
- MFA status
- encryption status
- immutability status
- offsite status
- audit status
- certificate status
- agent security
- update status
- ransomware risk
Every finding should contain:
- severity
- explanation
- affected object
- recommendation
---
# 18. Phase 16 — Ransomware Heuristics
Implement initial statistical detection.
Baseline:
- changed bytes
- changed file count
- extension distribution
- entropy indicators
- deletion rate
- unusual backup size
Do not make destructive decisions automatically.
Alert first.
---
# 19. Phase 17 — Reports
Implement:
- daily backup report
- weekly report
- monthly report
- failed backup report
- repository report
- recovery report
- security report
- compliance-oriented report
- RPO/RTO report
Export:
- CSV
- JSON
- PDF where practical
---
# 20. Phase 18 — Disaster Recovery
Implement and test:
## Scenario A
Control server lost.
Expected:
```text
Install Syncova
-> Restore configuration
-> Attach repository
-> Rebuild catalog
-> Restore VM
```
## Scenario B
Database lost.
Expected:
- restore database
- verify repository
- recover jobs/catalog
## Scenario C
Repository catalog lost.
Expected:
- scan repository
- reconstruct metadata
## Scenario D
Network interruption.
Expected:
- checkpoint
- retry
- resume
---
# 21. Phase 19 — Hardening
Perform:
- dependency scanning
- static analysis
- secret scanning
- API security tests
- RBAC tests
- authentication tests
- path traversal tests
- injection tests
- SSRF tests
- XSS tests
- privilege escalation tests
Harden:
- headers
- cookies
- CORS
- rate limits
- TLS
- service permissions
- filesystem permissions
---
# 22. Phase 20 — Performance Testing
Test:
- 1 large VM
- many small files
- multiple concurrent jobs
- high throughput repository
- slow repository
- slow network
- CPU-constrained host
- memory-constrained host
Measure:
- throughput
- latency
- CPU
- memory
- disk IOPS
- network
- repository contention
---
# 23. Phase 21 — Chaos Testing
Simulate:
- process kill
- service restart
- network loss
- repository unavailable
- database unavailable
- disk full
- corrupted chunk
- corrupted manifest
- power interruption simulation
Every failure must produce a controlled result.
---
# 24. Phase 22 — Release Candidate
Freeze:
- API contracts
- backup format version
- migration process
- repository protocol
Perform:
- full E2E test
- upgrade test
- downgrade/rollback test where supported
- disaster recovery test
- security review
- performance review
---
# 25. Phase 23 — V1 Release
Release package must include:
- backend
- frontend
- Windows agent
- Linux agent
- repository service
- migration files
- documentation
- installation instructions
- recovery documentation
- security guide
- API docs
- troubleshooting guide
---
# 26. Acceptance Test Matrix
## Backup
- [ ] Windows file backup
- [ ] Linux file backup
- [ ] Proxmox VM backup
- [ ] incremental
- [ ] deduplication
- [ ] compression
- [ ] encryption
- [ ] retry
- [ ] resume
- [ ] bandwidth control
## Recovery
- [ ] file restore
- [ ] folder restore
- [ ] Proxmox restore
- [ ] alternate-host restore
- [ ] restore validation
- [ ] interrupted restore recovery
## Repository
- [ ] local repository
- [ ] hardened repository
- [ ] integrity scan
- [ ] corruption detection
- [ ] catalog rebuild
- [ ] immutable retention
## Security
- [ ] RBAC
- [ ] MFA
- [ ] audit
- [ ] TLS
- [ ] secret protection
- [ ] privilege boundaries
## Monitoring
- [ ] dashboard
- [ ] metrics
- [ ] graphs
- [ ] alerts
- [ ] capacity forecast
- [ ] health
## Reliability
- [ ] control server restart
- [ ] database restart
- [ ] repository restart
- [ ] agent restart
- [ ] network interruption
- [ ] corrupted data
---
# 27. Recommended First Engineering Sprint
Do NOT begin with the dashboard.
The first production-quality vertical slice should be:
```text
PostgreSQL
|
Control API
|
Repository
|
Backup Engine
|
Test Source
|
Backup
|
Manifest
|
Integrity Verification
|
Restore
```
Once this works reliably, add agents and Proxmox.
The backup/recovery engine is the product's foundation.
---
# 28. Development Rules
1. Never fake a completed backup.
2. Never hide integrity failures.
3. Never store backup payloads in PostgreSQL.
4. Never put secrets in logs.
5. Never trust frontend authorization.
6. Never make destructive actions silent.
7. Never load complete backups into RAM.
8. Never make Proxmox code part of the generic backup engine.
9. Never mark a partial backup successful.
10. Never release a backup feature without a recovery test.
---
# 29. V1 Success Definition
Syncova V1 succeeds when a small IT team can:
1. Install Syncova.
2. Add a repository.
3. Add a Proxmox cluster.
4. Discover VMs.
5. Create a backup job.
6. Run a backup.
7. See live progress.
8. Verify the backup.
9. See meaningful statistics.
10. Restore the VM.
11. Confirm that it boots.
12. Understand the security/recovery state from the dashboard.
The experience should feel simple even though the underlying system is sophisticated.