Enterprise-Backup-, Recovery-, Verification-, Security- und Monitoring-Plattform fuer Proxmox VE, Windows, Linux und Dateisysteme. Der Leitsatz, der fast jede Entscheidung erklaert: Ein Backup gilt erst als vertrauenswuerdig, wenn Integritaet geprueft und Wiederherstellbarkeit nachgewiesen wurde. Deshalb steigt ein Wiederherstellungspunkt erst nach einem tatsaechlich durchgefuehrten Restore-Test auf "recoverable", und Unbekanntes geht in keine Bewertung als "gut" ein. Umfang (Phasen 0-23): - Repository Engine: inhaltsadressierte Bloecke, atomares Commit-Protokoll, Katalogaufbau allein aus den Manifesten — ohne Datenbank - Backup Engine: inhaltsabhaengiges Chunking, Deduplizierung trotz Verschluesselung, zstd, AES-256-GCM, Streaming mit Gegendruck - Agenten fuer Windows und Linux mit Auftragsabholung (Pull-Modell) - Proxmox-Provider mit beiden Zugriffswegen auf die Sicherungsarchive - Scheduler, Recovery Engine mit Pruefpunkt, Verification, Unveraenderlichkeit - Weboberflaeche, Kennzahlen, Meldungen, Berichte, Security Center, Ransomware-Heuristik (meldet, handelt nie) - Disaster Recovery, Haertung, Leistungsmessung, Chaos Testing - Eingefrorene Vertraege fuer API, Migrationen, Backup-Format und Repository - Auslieferungspaket fuer linux/amd64, linux/arm64 und windows/amd64 Nicht enthalten und als solches gekennzeichnet: Kapazitaetsprognose, Backup Copy, Changed Block Tracking bei Proxmox, erweiterte Attribute und ACLs. Gebaut, aber nie auf echter Hardware gefahren: der Windows-Dienst, die systemd-Einheit und der verpflichtende Proxmox-Meilenstein — ob eine wiederhergestellte VM startet, ist ungeprueft. Einzelheiten in CHANGELOG.md und docs/release-candidate.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
836 lines
12 KiB
Markdown
836 lines
12 KiB
Markdown
# Syncova Backups V1 — Implementation Plan
|
|
|
|
## 1. Objective
|
|
|
|
Build Syncova V1 as a production-oriented backup and recovery platform with a staged implementation process.
|
|
|
|
Do not attempt to implement all functionality simultaneously.
|
|
|
|
Each phase ends with:
|
|
|
|
- working software
|
|
- automated tests
|
|
- documentation
|
|
- security review
|
|
- measurable acceptance criteria
|
|
|
|
## 2. Phase 0 — Product Foundation
|
|
|
|
### Deliverables
|
|
|
|
- repository structure
|
|
- coding standards
|
|
- architecture documentation
|
|
- CI pipeline
|
|
- dependency management
|
|
- local development environment
|
|
- PostgreSQL
|
|
- migration framework
|
|
- basic frontend shell
|
|
- backend health endpoint
|
|
|
|
### Suggested repository
|
|
|
|
```text
|
|
syncova/
|
|
apps/
|
|
api/
|
|
web/
|
|
agent/
|
|
worker/
|
|
packages/
|
|
backup/
|
|
repository/
|
|
crypto/
|
|
providers/
|
|
metrics/
|
|
auth/
|
|
migrations/
|
|
docs/
|
|
tests/
|
|
deployment/
|
|
```
|
|
|
|
### Exit Criteria
|
|
|
|
- Backend starts.
|
|
- Frontend starts.
|
|
- Database connects.
|
|
- Migration runs.
|
|
- CI executes tests.
|
|
- No hardcoded credentials.
|
|
|
|
---
|
|
|
|
# 3. Phase 1 — Identity and Security Foundation
|
|
|
|
Implement:
|
|
|
|
- users
|
|
- roles
|
|
- permissions
|
|
- login
|
|
- sessions/tokens
|
|
- password hashing
|
|
- TOTP MFA
|
|
- audit logging
|
|
- TLS configuration
|
|
- secret abstraction
|
|
|
|
### Exit Criteria
|
|
|
|
- Admin can log in.
|
|
- MFA works.
|
|
- Unauthorized users cannot access protected APIs.
|
|
- Privileged operations are audited.
|
|
- Passwords are never stored plaintext.
|
|
|
|
---
|
|
|
|
# 4. Phase 2 — Repository Engine
|
|
|
|
This is the first major engineering milestone.
|
|
|
|
Implement:
|
|
|
|
- repository abstraction
|
|
- local filesystem repository
|
|
- hardened repository mode
|
|
- chunk storage
|
|
- manifests
|
|
- atomic commit
|
|
- repository catalog
|
|
- integrity metadata
|
|
- repository health
|
|
- repository scan
|
|
|
|
### Tests
|
|
|
|
- write chunk
|
|
- read chunk
|
|
- detect corruption
|
|
- interrupted write
|
|
- restart
|
|
- catalog rebuild
|
|
|
|
### Exit Criteria
|
|
|
|
A repository can be created, written, scanned, restarted, and rebuilt without PostgreSQL containing the only copy of backup metadata.
|
|
|
|
---
|
|
|
|
# 5. Phase 3 — Backup Format
|
|
|
|
Implement versioned Syncova format.
|
|
|
|
Must contain:
|
|
|
|
- header
|
|
- source metadata
|
|
- backup metadata
|
|
- manifest
|
|
- chunk references
|
|
- integrity information
|
|
- encryption metadata
|
|
- completion marker
|
|
|
|
### Tests
|
|
|
|
- serialize
|
|
- deserialize
|
|
- version compatibility
|
|
- invalid manifest
|
|
- missing chunk
|
|
- corrupted chunk
|
|
- incomplete backup
|
|
|
|
---
|
|
|
|
# 6. Phase 4 — Backup Engine Core
|
|
|
|
Implement streaming pipeline:
|
|
|
|
```text
|
|
Read
|
|
-> Chunk
|
|
-> Hash
|
|
-> Dedup
|
|
-> Compress
|
|
-> Encrypt
|
|
-> Write
|
|
-> Manifest
|
|
-> Commit
|
|
```
|
|
|
|
Implement:
|
|
|
|
- bounded worker pools
|
|
- backpressure
|
|
- checkpoints
|
|
- cancellation
|
|
- retry
|
|
- progress reporting
|
|
|
|
### Metrics
|
|
|
|
Collect:
|
|
|
|
- bytes processed
|
|
- bytes written
|
|
- throughput
|
|
- compression ratio
|
|
- dedup ratio
|
|
- duration
|
|
|
|
### Exit Criteria
|
|
|
|
A large test dataset can be backed up without loading the dataset into memory.
|
|
|
|
---
|
|
|
|
# 7. Phase 5 — Windows Agent
|
|
|
|
Implement:
|
|
|
|
- Windows service
|
|
- enrollment
|
|
- heartbeat
|
|
- file discovery
|
|
- file backup
|
|
- restore
|
|
- secure communication
|
|
- local logging
|
|
|
|
### Tests
|
|
|
|
- service restart
|
|
- network interruption
|
|
- permission errors
|
|
- large files
|
|
- many small files
|
|
- incremental backup
|
|
|
|
### Exit Criteria
|
|
|
|
A Windows system can be backed up and restored reliably.
|
|
|
|
---
|
|
|
|
# 8. Phase 6 — Linux Agent
|
|
|
|
Implement:
|
|
|
|
- systemd service
|
|
- enrollment
|
|
- heartbeat
|
|
- filesystem backup
|
|
- restore
|
|
- secure communication
|
|
|
|
Test:
|
|
|
|
- permissions
|
|
- symlinks
|
|
- special files as applicable
|
|
- interrupted jobs
|
|
- incremental operation
|
|
|
|
---
|
|
|
|
# 9. Phase 7 — Proxmox Provider
|
|
|
|
Implement provider interface and Proxmox implementation.
|
|
|
|
### Discovery
|
|
|
|
- cluster
|
|
- hosts
|
|
- VMs
|
|
- VM storage
|
|
- configuration
|
|
|
|
### Backup
|
|
|
|
- VM metadata
|
|
- VM disks
|
|
- snapshots where appropriate
|
|
- changed blocks where available
|
|
- consistency state
|
|
|
|
### Restore
|
|
|
|
- original host
|
|
- alternate host
|
|
- new VM
|
|
|
|
### Exit Criteria
|
|
|
|
A test Proxmox environment can:
|
|
|
|
```text
|
|
Discover VM
|
|
-> Backup VM
|
|
-> Verify
|
|
-> Delete test VM
|
|
-> Restore VM
|
|
-> Boot VM
|
|
-> Validate VM
|
|
```
|
|
|
|
This is a mandatory end-to-end milestone.
|
|
|
|
---
|
|
|
|
# 10. Phase 8 — Scheduler and Job Management
|
|
|
|
Implement:
|
|
|
|
- jobs
|
|
- schedules
|
|
- priorities
|
|
- concurrency
|
|
- retries
|
|
- backoff
|
|
- job dependencies
|
|
- bandwidth limits
|
|
- maintenance windows
|
|
|
|
### UI
|
|
|
|
Create the backup wizard:
|
|
|
|
1. Name
|
|
2. Source
|
|
3. Schedule
|
|
4. Repository
|
|
5. Retention
|
|
6. Security
|
|
7. Verification
|
|
8. Notifications
|
|
9. Review
|
|
10. Create
|
|
|
|
---
|
|
|
|
# 11. Phase 9 — Recovery Engine
|
|
|
|
Implement:
|
|
|
|
- file restore
|
|
- folder restore
|
|
- disk restore
|
|
- full system restore where supported
|
|
- Proxmox VM restore
|
|
- restore validation
|
|
- restore sessions
|
|
- checkpoints
|
|
- progress
|
|
|
|
### Safety
|
|
|
|
Restores to production must require explicit confirmation.
|
|
|
|
---
|
|
|
|
# 12. Phase 10 — Verification and Recovery Assurance
|
|
|
|
Implement:
|
|
|
|
- chunk verification
|
|
- manifest verification
|
|
- chain validation
|
|
- repository integrity scans
|
|
- restore-point validation
|
|
- automated recovery tests
|
|
- Recovery Assurance score
|
|
|
|
### Score Inputs
|
|
|
|
- backup freshness
|
|
- verification
|
|
- restore test
|
|
- RPO
|
|
- RTO
|
|
- immutability
|
|
- offsite
|
|
- encryption
|
|
- repository health
|
|
- anomalies
|
|
|
|
### Exit Criteria
|
|
|
|
A backup can be objectively classified as:
|
|
|
|
- successful
|
|
- verified
|
|
- recoverable
|
|
- failed
|
|
- corrupted
|
|
|
|
---
|
|
|
|
# 13. Phase 11 — Immutability
|
|
|
|
Implement hardened repository protections.
|
|
|
|
Requirements:
|
|
|
|
- retention lock
|
|
- immutable-until metadata
|
|
- restricted delete
|
|
- audit
|
|
- cleanup protection
|
|
|
|
Where the storage layer supports stronger WORM/object-lock behavior, integrate it.
|
|
|
|
Do not claim true immutability if the underlying storage cannot enforce it.
|
|
|
|
---
|
|
|
|
# 14. Phase 12 — Web Dashboard
|
|
|
|
Build the main UI.
|
|
|
|
### Pages
|
|
|
|
- Dashboard
|
|
- Backup Jobs
|
|
- Protected Systems
|
|
- Recovery
|
|
- Recovery Points
|
|
- Repositories
|
|
- Proxmox
|
|
- Agents
|
|
- Alerts
|
|
- Events
|
|
- Security Center
|
|
- Reports
|
|
- Users
|
|
- Roles
|
|
- Settings
|
|
|
|
### Dashboard widgets
|
|
|
|
- protected systems
|
|
- backup success rate
|
|
- failed jobs
|
|
- critical alerts
|
|
- storage
|
|
- capacity forecast
|
|
- Recovery Assurance
|
|
- Security Score
|
|
- RPO compliance
|
|
- repository health
|
|
|
|
---
|
|
|
|
# 15. Phase 13 — Metrics and Graphs
|
|
|
|
Implement metrics storage and chart APIs.
|
|
|
|
Required graphs:
|
|
|
|
- backup success/failure
|
|
- duration
|
|
- throughput
|
|
- processed/written data
|
|
- storage growth
|
|
- deduplication
|
|
- compression
|
|
- repository capacity
|
|
- agent resources
|
|
- RPO
|
|
- recovery duration
|
|
- verification results
|
|
|
|
Time ranges:
|
|
|
|
- 1h
|
|
- 24h
|
|
- 7d
|
|
- 30d
|
|
- 90d
|
|
- 1y
|
|
- custom
|
|
|
|
---
|
|
|
|
# 16. Phase 14 — Alerts and Notifications
|
|
|
|
Implement:
|
|
|
|
- alert rules
|
|
- severity
|
|
- acknowledge
|
|
- resolve
|
|
- email
|
|
- webhook
|
|
|
|
Initial rules:
|
|
|
|
- backup failure
|
|
- repeated backup failure
|
|
- repository unavailable
|
|
- repository >80%
|
|
- repository >90%
|
|
- agent offline
|
|
- RPO violation
|
|
- verification failure
|
|
- certificate expiry
|
|
- possible ransomware activity
|
|
- immutable copy missing
|
|
|
|
---
|
|
|
|
# 17. Phase 15 — Security Center
|
|
|
|
Implement:
|
|
|
|
- Security Score
|
|
- MFA status
|
|
- encryption status
|
|
- immutability status
|
|
- offsite status
|
|
- audit status
|
|
- certificate status
|
|
- agent security
|
|
- update status
|
|
- ransomware risk
|
|
|
|
Every finding should contain:
|
|
|
|
- severity
|
|
- explanation
|
|
- affected object
|
|
- recommendation
|
|
|
|
---
|
|
|
|
# 18. Phase 16 — Ransomware Heuristics
|
|
|
|
Implement initial statistical detection.
|
|
|
|
Baseline:
|
|
|
|
- changed bytes
|
|
- changed file count
|
|
- extension distribution
|
|
- entropy indicators
|
|
- deletion rate
|
|
- unusual backup size
|
|
|
|
Do not make destructive decisions automatically.
|
|
|
|
Alert first.
|
|
|
|
---
|
|
|
|
# 19. Phase 17 — Reports
|
|
|
|
Implement:
|
|
|
|
- daily backup report
|
|
- weekly report
|
|
- monthly report
|
|
- failed backup report
|
|
- repository report
|
|
- recovery report
|
|
- security report
|
|
- compliance-oriented report
|
|
- RPO/RTO report
|
|
|
|
Export:
|
|
|
|
- CSV
|
|
- JSON
|
|
- PDF where practical
|
|
|
|
---
|
|
|
|
# 20. Phase 18 — Disaster Recovery
|
|
|
|
Implement and test:
|
|
|
|
## Scenario A
|
|
|
|
Control server lost.
|
|
|
|
Expected:
|
|
|
|
```text
|
|
Install Syncova
|
|
-> Restore configuration
|
|
-> Attach repository
|
|
-> Rebuild catalog
|
|
-> Restore VM
|
|
```
|
|
|
|
## Scenario B
|
|
|
|
Database lost.
|
|
|
|
Expected:
|
|
|
|
- restore database
|
|
- verify repository
|
|
- recover jobs/catalog
|
|
|
|
## Scenario C
|
|
|
|
Repository catalog lost.
|
|
|
|
Expected:
|
|
|
|
- scan repository
|
|
- reconstruct metadata
|
|
|
|
## Scenario D
|
|
|
|
Network interruption.
|
|
|
|
Expected:
|
|
|
|
- checkpoint
|
|
- retry
|
|
- resume
|
|
|
|
---
|
|
|
|
# 21. Phase 19 — Hardening
|
|
|
|
Perform:
|
|
|
|
- dependency scanning
|
|
- static analysis
|
|
- secret scanning
|
|
- API security tests
|
|
- RBAC tests
|
|
- authentication tests
|
|
- path traversal tests
|
|
- injection tests
|
|
- SSRF tests
|
|
- XSS tests
|
|
- privilege escalation tests
|
|
|
|
Harden:
|
|
|
|
- headers
|
|
- cookies
|
|
- CORS
|
|
- rate limits
|
|
- TLS
|
|
- service permissions
|
|
- filesystem permissions
|
|
|
|
---
|
|
|
|
# 22. Phase 20 — Performance Testing
|
|
|
|
Test:
|
|
|
|
- 1 large VM
|
|
- many small files
|
|
- multiple concurrent jobs
|
|
- high throughput repository
|
|
- slow repository
|
|
- slow network
|
|
- CPU-constrained host
|
|
- memory-constrained host
|
|
|
|
Measure:
|
|
|
|
- throughput
|
|
- latency
|
|
- CPU
|
|
- memory
|
|
- disk IOPS
|
|
- network
|
|
- repository contention
|
|
|
|
---
|
|
|
|
# 23. Phase 21 — Chaos Testing
|
|
|
|
Simulate:
|
|
|
|
- process kill
|
|
- service restart
|
|
- network loss
|
|
- repository unavailable
|
|
- database unavailable
|
|
- disk full
|
|
- corrupted chunk
|
|
- corrupted manifest
|
|
- power interruption simulation
|
|
|
|
Every failure must produce a controlled result.
|
|
|
|
---
|
|
|
|
# 24. Phase 22 — Release Candidate
|
|
|
|
Freeze:
|
|
|
|
- API contracts
|
|
- backup format version
|
|
- migration process
|
|
- repository protocol
|
|
|
|
Perform:
|
|
|
|
- full E2E test
|
|
- upgrade test
|
|
- downgrade/rollback test where supported
|
|
- disaster recovery test
|
|
- security review
|
|
- performance review
|
|
|
|
---
|
|
|
|
# 25. Phase 23 — V1 Release
|
|
|
|
Release package must include:
|
|
|
|
- backend
|
|
- frontend
|
|
- Windows agent
|
|
- Linux agent
|
|
- repository service
|
|
- migration files
|
|
- documentation
|
|
- installation instructions
|
|
- recovery documentation
|
|
- security guide
|
|
- API docs
|
|
- troubleshooting guide
|
|
|
|
---
|
|
|
|
# 26. Acceptance Test Matrix
|
|
|
|
## Backup
|
|
|
|
- [ ] Windows file backup
|
|
- [ ] Linux file backup
|
|
- [ ] Proxmox VM backup
|
|
- [ ] incremental
|
|
- [ ] deduplication
|
|
- [ ] compression
|
|
- [ ] encryption
|
|
- [ ] retry
|
|
- [ ] resume
|
|
- [ ] bandwidth control
|
|
|
|
## Recovery
|
|
|
|
- [ ] file restore
|
|
- [ ] folder restore
|
|
- [ ] Proxmox restore
|
|
- [ ] alternate-host restore
|
|
- [ ] restore validation
|
|
- [ ] interrupted restore recovery
|
|
|
|
## Repository
|
|
|
|
- [ ] local repository
|
|
- [ ] hardened repository
|
|
- [ ] integrity scan
|
|
- [ ] corruption detection
|
|
- [ ] catalog rebuild
|
|
- [ ] immutable retention
|
|
|
|
## Security
|
|
|
|
- [ ] RBAC
|
|
- [ ] MFA
|
|
- [ ] audit
|
|
- [ ] TLS
|
|
- [ ] secret protection
|
|
- [ ] privilege boundaries
|
|
|
|
## Monitoring
|
|
|
|
- [ ] dashboard
|
|
- [ ] metrics
|
|
- [ ] graphs
|
|
- [ ] alerts
|
|
- [ ] capacity forecast
|
|
- [ ] health
|
|
|
|
## Reliability
|
|
|
|
- [ ] control server restart
|
|
- [ ] database restart
|
|
- [ ] repository restart
|
|
- [ ] agent restart
|
|
- [ ] network interruption
|
|
- [ ] corrupted data
|
|
|
|
---
|
|
|
|
# 27. Recommended First Engineering Sprint
|
|
|
|
Do NOT begin with the dashboard.
|
|
|
|
The first production-quality vertical slice should be:
|
|
|
|
```text
|
|
PostgreSQL
|
|
|
|
|
Control API
|
|
|
|
|
Repository
|
|
|
|
|
Backup Engine
|
|
|
|
|
Test Source
|
|
|
|
|
Backup
|
|
|
|
|
Manifest
|
|
|
|
|
Integrity Verification
|
|
|
|
|
Restore
|
|
```
|
|
|
|
Once this works reliably, add agents and Proxmox.
|
|
|
|
The backup/recovery engine is the product's foundation.
|
|
|
|
---
|
|
|
|
# 28. Development Rules
|
|
|
|
1. Never fake a completed backup.
|
|
2. Never hide integrity failures.
|
|
3. Never store backup payloads in PostgreSQL.
|
|
4. Never put secrets in logs.
|
|
5. Never trust frontend authorization.
|
|
6. Never make destructive actions silent.
|
|
7. Never load complete backups into RAM.
|
|
8. Never make Proxmox code part of the generic backup engine.
|
|
9. Never mark a partial backup successful.
|
|
10. Never release a backup feature without a recovery test.
|
|
|
|
---
|
|
|
|
# 29. V1 Success Definition
|
|
|
|
Syncova V1 succeeds when a small IT team can:
|
|
|
|
1. Install Syncova.
|
|
2. Add a repository.
|
|
3. Add a Proxmox cluster.
|
|
4. Discover VMs.
|
|
5. Create a backup job.
|
|
6. Run a backup.
|
|
7. See live progress.
|
|
8. Verify the backup.
|
|
9. See meaningful statistics.
|
|
10. Restore the VM.
|
|
11. Confirm that it boots.
|
|
12. Understand the security/recovery state from the dashboard.
|
|
|
|
The experience should feel simple even though the underlying system is sophisticated.
|