Enterprise-Backup-, Recovery-, Verification-, Security- und Monitoring-Plattform fuer Proxmox VE, Windows, Linux und Dateisysteme. Der Leitsatz, der fast jede Entscheidung erklaert: Ein Backup gilt erst als vertrauenswuerdig, wenn Integritaet geprueft und Wiederherstellbarkeit nachgewiesen wurde. Deshalb steigt ein Wiederherstellungspunkt erst nach einem tatsaechlich durchgefuehrten Restore-Test auf "recoverable", und Unbekanntes geht in keine Bewertung als "gut" ein. Umfang (Phasen 0-23): - Repository Engine: inhaltsadressierte Bloecke, atomares Commit-Protokoll, Katalogaufbau allein aus den Manifesten — ohne Datenbank - Backup Engine: inhaltsabhaengiges Chunking, Deduplizierung trotz Verschluesselung, zstd, AES-256-GCM, Streaming mit Gegendruck - Agenten fuer Windows und Linux mit Auftragsabholung (Pull-Modell) - Proxmox-Provider mit beiden Zugriffswegen auf die Sicherungsarchive - Scheduler, Recovery Engine mit Pruefpunkt, Verification, Unveraenderlichkeit - Weboberflaeche, Kennzahlen, Meldungen, Berichte, Security Center, Ransomware-Heuristik (meldet, handelt nie) - Disaster Recovery, Haertung, Leistungsmessung, Chaos Testing - Eingefrorene Vertraege fuer API, Migrationen, Backup-Format und Repository - Auslieferungspaket fuer linux/amd64, linux/arm64 und windows/amd64 Nicht enthalten und als solches gekennzeichnet: Kapazitaetsprognose, Backup Copy, Changed Block Tracking bei Proxmox, erweiterte Attribute und ACLs. Gebaut, aber nie auf echter Hardware gefahren: der Windows-Dienst, die systemd-Einheit und der verpflichtende Proxmox-Meilenstein — ob eine wiederhergestellte VM startet, ist ungeprueft. Einzelheiten in CHANGELOG.md und docs/release-candidate.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
9.0 KiB
Syncova Backups V1 — System Architecture
1. Purpose
Syncova is an enterprise-grade backup, recovery, verification, security, and monitoring platform. V1 focuses on:
- Proxmox VE
- Windows Server / Windows clients
- Linux systems
- Physical systems
- Files and folders
VMware is explicitly out of scope for V1, but the architecture must support future provider implementations.
Core principle:
A backup is not considered trustworthy until integrity and recoverability are proven.
2. High-Level Architecture
+----------------------+
| Web UI |
| React + TypeScript |
+----------+-----------+
|
HTTPS / WebSocket
|
+----------v-----------+
| API Gateway |
| REST / Auth / RBAC |
+----------+-----------+
|
+---------------+----------------+
| |
+---------v---------+ +--------v--------+
| Control Plane | | Monitoring |
| Jobs / Policies | | Metrics / Alerts|
+---------+---------+ +--------+--------+
| |
+---------------+----------------+
|
+----------v-----------+
| Backup Orchestrator |
+----------+-----------+
|
+---------------------+----------------------+
| | |
+------v------+ +------v------+ +------v------+
| Proxmox | | Windows | | Linux |
| Provider | | Agent | | Agent |
+------+------+ +------+------+ +------+------+
| | |
+---------------------+----------------------+
|
+----------v-----------+
| Backup Engine |
| chunk/dedup/compress |
| encrypt/integrity |
+----------+-----------+
|
+----------v-----------+
| Repository Service |
+----------+-----------+
|
+-------------------+-------------------+
| | |
Local/Hardened S3-compatible Secondary
Repository Storage Repository
3. Design Principles
- Security by default.
- Least privilege.
- Recovery-first design.
- Repository data independent of PostgreSQL.
- Control plane failure must not destroy backup data.
- Streaming processing; never load complete backups into RAM.
- Provider abstraction for hypervisors.
- Versioned backup format.
- Explicit states; no silent failures.
- Every destructive operation is auditable.
4. Services
4.1 API Service
Responsibilities:
- HTTP API
- authentication
- authorization
- request validation
- rate limiting
- API versioning
- audit integration
4.2 Control Service
Responsibilities:
- jobs
- policies
- sources
- repositories
- orchestration
- configuration
- scheduling
4.3 Scheduler
Responsibilities:
- schedules
- retry/backoff
- concurrency
- priorities
- backup windows
- job dependencies
4.4 Backup Engine
Pipeline:
Source Read
-> Change Detection
-> Chunking
-> Hash
-> Dedup Lookup
-> Compression
-> Encryption
-> Repository Write
-> Manifest Commit
-> Verification
Backpressure must be used between pipeline stages.
4.5 Repository Service
Responsibilities:
- object/chunk writes
- manifests
- atomic commits
- retention operations
- integrity scanning
- repository health
- repository rebuild
PostgreSQL is not the source of truth for backup payloads.
4.6 Verification Service
Responsibilities:
- integrity checks
- chain validation
- restore-point validation
- automated restore tests
- recovery assurance
4.7 Monitoring Service
Responsibilities:
- metrics
- events
- alerts
- health checks
- anomaly detection
- capacity forecasts
5. Provider Architecture
VirtualizationProvider
|
+-- ProxmoxProvider (V1)
+-- VMwareProvider (future)
+-- HyperVProvider (future)
Generic interface:
Connect
Disconnect
ListClusters
ListHosts
ListVMs
GetVMInfo
GetVMDisks
GetVMMetaData
CreateSnapshot
RemoveSnapshot
ReadChangedBlocks
RestoreVM
Provider-specific logic must not leak into the backup engine.
6. Agent Architecture
Agents run as native services:
- Windows Service
- Linux systemd service
Responsibilities:
- secure registration
- heartbeat
- source discovery
- file/system reads
- data streaming
- restore execution
- local health reporting
Agents must use minimal privileges.
7. Security Architecture
Authentication
V1:
- username/password
- TOTP MFA
Architecture-ready:
- WebAuthn/passkeys
- OIDC
- SAML
- LDAP/AD
- Entra ID
Authorization
RBAC with explicit permissions.
Critical actions may require step-up MFA or future four-eyes approval.
Transport
TLS for all remote communication.
Secrets
Never store or log secrets in plaintext. Use encrypted secret storage and a key hierarchy.
8. Backup Format
A backup consists of:
Header
Format Version
Backup ID
Source ID
Creation Time
Encryption Metadata
Manifest
Files/Disks
Chunk References
Parent Backup
Consistency Level
Chunk Index
Chunk ID
Offset
Length
Integrity Metadata
Encrypted Chunks
Footer
Manifest Hash
Completion Marker
The format must support version negotiation and future extensions.
9. Repository Layout
Logical layout:
repository/
format/
manifests/
chunks/
indexes/
journals/
verification/
metadata/
Exact physical layout may evolve, but repositories must remain self-describing and rebuildable.
10. Repository Commit Protocol
Use staged writes:
- Create session.
- Write chunks.
- Write manifest.
- Verify manifest.
- Atomically commit completion marker.
- Update catalog.
- Publish successful backup state.
A backup without a valid completion marker is incomplete.
11. Recovery Architecture
Recovery must work even if the control server is rebuilt.
Flow:
Attach Repository
-> Discover Format
-> Scan Manifests
-> Validate Chains
-> Rebuild Catalog
-> Select Restore Point
-> Validate Dependencies
-> Restore
-> Verify
12. Proxmox Recovery
Support:
- original host
- alternate host
- restore as new VM
- VM configuration restoration
- disk restoration
- metadata restoration
Use official Proxmox APIs where possible.
13. Metrics
Core metrics include:
- backup.duration
- backup.bytes_processed
- backup.bytes_written
- backup.throughput
- backup.success
- backup.failure
- repository.capacity
- repository.used
- repository.free
- repository.latency
- dedup.ratio
- compression.ratio
- verification.duration
- recovery.duration
- agent.cpu
- agent.memory
- network.throughput
Metric retention should use rollups.
14. Recovery Assurance
Each protected system gets a score based on:
- recent successful backup
- verification
- recovery test
- RPO compliance
- RTO compliance
- immutability
- offsite copy
- encryption
- repository health
- anomaly/ransomware risk
The score must always be explainable.
15. Failure Model
Failures must be classified as:
- transient
- permanent
- integrity
- authentication
- authorization
- repository
- source
- network
- configuration
- security
Retries use bounded exponential backoff.
16. Scaling
V1 should support small environments but have a path to 1000+ protected systems.
Avoid shared global locks. Use:
- worker pools
- bounded concurrency
- partitionable queues
- indexed queries
- streaming
- asynchronous jobs
17. Deployment
Initial supported deployment:
- Linux server
- VM
- dedicated backup appliance
Components should be packaged so that later containerized deployment is possible.
18. Observability
Every request/job/session has:
- correlation ID
- job ID
- task ID
- structured logs
Expose:
- liveness
- readiness
- metrics
- health checks
19. Disaster Recovery
Back up Syncova configuration and provide:
- control-server rebuild
- database restore
- repository attach
- catalog rebuild
- key recovery
20. Non-Goals for V1
Do not implement:
- VMware
- Kubernetes
- Microsoft 365
- Exchange application-aware backup
- broad SaaS backup
- complex multi-tenant MSP architecture
- AI-dependent functionality
The architecture must leave room for these later.