Storage Platform

Software-defined storage & resilience

Data is your most valuable asset.
Hardware failure is not a matter of if, but when.

In short: whitesky’s storage uses erasure coding across independent storage blocks (3-6 servers each) with automatic self-healing and rebuild. This tolerates multiple disk/server failures without data loss. Depending on the erasure-coding policy, storage overhead can be as low as ~33% — far below the 200% of triple replication.

Already invested in a SAN? whitesky also runs virtual machines on your existing external storage arrays. See External storage arrays below.

The difference between data loss and data safety is not marketing claims or SLAs — it is architectural decisions made from day one. The whitesky storage platform is engineered for reality: disks fail, servers fail, networks partition, and maintenance must happen without downtime.

This page explains how whitesky delivers data safety by design, from high-level principles to the technical foundations underneath.


Engineering for reality, not best-case scenarios

Traditional storage platforms are often optimized for benchmarks and ideal conditions. whitesky starts from a different assumption: failure is expected.

Every architectural choice is shaped by this premise:

  • Safety takes precedence over raw efficiency
  • Scalability is built in, not retrofitted
  • Failure is isolated, not amplified
  • Recovery is automatic, not heroic

Architecture determines resilience. Hardware fails. Power fluctuates. Networks partition.
Your storage platform must handle these realities transparently.


whitesky storage at a glance

The whitesky storage platform is a distributed, software-defined system designed to withstand cascading failures that plague traditional infrastructure.

Even during:

  • disk failures
  • server outages
  • network partitions
  • planned maintenance

the platform maintains data availability and consistency.

Platform layers

  • Compute layer: Virtual machine orchestration and workload management with seamless failover.
  • Block storage layer: High-performance virtual disks with distributed transaction logging and cheap snapshotting.
  • Object storage backend: Erasure-coded data distribution across fault domains with automatic self-healing.
  • Backup layer: Independent snapshot architecture using S3-compatible storage. With S3 object locking (governance or compliance mode), backups cannot be altered or deleted through the whitesky or S3 APIs.
  • External storage (optional): Virtual disks on your existing SAN arrays, managed through the same portal and API.
Block storage: N+2 optionsTwo built-in block storage options, physical NVMe storage attached directly to a VM and the whitesky software-defined storage, plus any number of external SAN arrays over Fibre Channel or Ethernet, with compatibility certified per vendor and SAN type. Each virtual disk of a VM uses the option that fits it.Built in: 2 optionsPlus NVMNVMePhysical storageLocal NVMe, attacheddirectly to a VMMaximum performancewhitesky SDSSoftware-defined, erasure-coded and self-healingDefault for new locations× NExternal SANCertified per SAN type,over FC or EthernetKeep your arraysEach virtual disk of a VM uses the option that fits it
Block storage: N+2 optionsTwo built-in block storage options, physical NVMe storage attached directly to a VM and the whitesky software-defined storage, plus any number of external SAN arrays over Fibre Channel or Ethernet, with compatibility certified per vendor and SAN type. Each virtual disk of a VM uses the option that fits it.Built in: 2 optionsPhysical storageLocal NVMe, attacheddirectly to a VMMaximum performanceNVMewhitesky SDSSoftware-defined, erasure-coded and self-healingDefault for new locationsPlus NExternal SAN × NCertified per SAN type,over FC or EthernetKeep your arraysEach virtual disk of a VM usesthe option that fits it
Block storage, N+2: two options built in, plus any number of external SAN arrays.

Multiple layers of fault tolerance

Device-level protection: beyond RAID

Traditional RAID introduces single points of failure. whitesky uses erasure coding instead.

Data is split into fragments with calculated redundancy. These fragments are distributed across different physical disks.
Even multiple simultaneous disk failures do not result in data loss.

Erasure coding with a 12+4 policyAn object is split into 12 data fragments and 4 parity fragments, spread over disks in four servers. Four disks fail at once, spread over three servers, and the object is still readable: any 12 of the 16 fragments rebuild it.1 object12 data fragments4 paritySpread over disks in different servers:Server 1disk 1failedfaileddisk 4Server 2disk 1disk 2disk 3disk 4Server 3faileddisk 2disk 3disk 4Server 4disk 1disk 2disk 3failed✓ 4 disks lost, data intact: any 12 of the 16 fragments rebuild the object
Erasure coding with a 12+4 policyAn object is split into 12 data fragments and 4 parity fragments, spread over disks in four servers. Four disks fail at once, spread over three servers, and the object is still readable: any 12 of the 16 fragments rebuild it.1 object12 data fragments4 paritySpread over disks in different servers:Server 1disk 1failedfaileddisk 4Server 2disk 1disk 2disk 3disk 4Server 3faileddisk 2disk 3disk 4Server 4disk 1disk 2disk 3failed✓ 4 disks lost, data intact:any 12 of the 16 fragments rebuild it
Example with a 12+4 policy: four disks fail at once, spread over three servers, and the object is still complete. The policy, and with it the number of failures tolerated, is chosen per deployment.

Server-level protection: distributed by design

Fragments are deliberately spread across different physical servers. No single server ever holds critical data.

When a server fails:

  • data remains accessible
  • fragments are automatically rebuilt
  • redistribution happens without manual intervention

For decision makers: losing disks or servers does not mean losing data.
Failure is routine — not catastrophic.


Storage blocks: independent failure domains

Instead of monolithic storage clusters, whitesky uses storage blocks.

A storage block is a small, independent failure domain:

  • 3 to 6 storage servers per block
  • always tolerates the loss of 1 full server
  • and multiple disk failures at the same time
Storage blocks are independent failure domainsA cloud location made of three storage blocks with 4, 5 and 3 servers. A server in block B fails: block B rebuilds and keeps serving data, blocks A and C are unaffected.One cloud locationStorage block AserverserverserverserverStorage block Bserverserverserver downserverserverStorage block Cserverserverserver✓ unaffected↻ rebuilds, data stays available✓ unaffected
Storage blocks are independent failure domainsA cloud location made of three storage blocks with 4, 5 and 3 servers. A server in block B fails: block B rebuilds and keeps serving data, blocks A and C are unaffected.One cloud locationStorage block A✓ unaffectedStorage block B↻ rebuilds, data stays availableStorage block C✓ unaffected
A server failure stays inside its storage block. The other blocks keep running as if nothing happened.

Why this matters

  • Failure containment: Failures stay inside one block. There is no cascading “blast radius” across your entire cloud.
  • Predictable recovery: Each block has known recovery characteristics. No surprises during incidents.
  • Autonomous operation: Each block operates independently. Issues in one block never impact the operational state of others.
  • Multiple storage blocks combine into a full cloud location, enabling scale without sacrificing safety.

Scaling without compromising safety

Linear scale-out architecture

Capacity and performance scale by adding storage blocks. No re-architecture. No migration events. No redesign.

You can:

  • start with a single block for edge or small deployments
  • grow to dozens of blocks for regional datacenters

Each block adds predictable capacity, performance, and fault tolerance.

Efficient storage economics

Storage overhead depends on the erasure-coding policy in use. A policy is expressed as data + parity fragments, and the overhead is the parity fragments divided by the data fragments:

  • A 12+4 policy tolerates 4 simultaneous failures at ~33% overhead.
  • An 8+4 policy tolerates 4 failures at 50% overhead.
  • A 4+2 policy tolerates 2 failures at 50% overhead.
Storage overhead per redundancy policyExtra raw capacity needed per unit of data: 12+4 erasure coding about 33 percent, tolerating 4 failures; 8+4 erasure coding 50 percent, 4 failures; 4+2 erasure coding 50 percent, 2 failures; triple replication 200 percent, 2 failures.Extra raw capacity per unit of datawhitesky erasure codingreplication0%100%200%12+4 erasure codingtolerates 4 failures12+4 erasure coding: ~33% overhead, tolerates 4 failures~33%8+4 erasure codingtolerates 4 failures8+4 erasure coding: 50% overhead, tolerates 4 failures50%4+2 erasure codingtolerates 2 failures4+2 erasure coding: 50% overhead, tolerates 2 failures50%Triple replicationtolerates 2 failuresTriple replication: 200% overhead, tolerates 2 failures200%
Storage overhead per redundancy policyExtra raw capacity needed per unit of data: 12+4 erasure coding about 33 percent, tolerating 4 failures; 8+4 erasure coding 50 percent, 4 failures; 4+2 erasure coding 50 percent, 2 failures; triple replication 200 percent, 2 failures.Extra raw capacity per unit of datawhitesky erasure codingreplication0%100%200%12+4 erasure codingtolerates 4 failures12+4 erasure coding: ~33% overhead, tolerates 4 failures~33%8+4 erasure codingtolerates 4 failures8+4 erasure coding: 50% overhead, tolerates 4 failures50%4+2 erasure codingtolerates 2 failures4+2 erasure coding: 50% overhead, tolerates 2 failures50%Triple replicationtolerates 2 failuresTriple replication: 200% overhead, tolerates 2 failures200%
Extra raw capacity needed per unit of data, per redundancy policy.

So overhead as low as ~33% is achievable with wider policies, while smaller blocks trade some efficiency for tighter failure domains. Every one of these is dramatically more efficient than triple replication, which fixes overhead at 200% of the data (3× raw capacity per usable TB). You choose the balance of resilience and efficiency — and even the most conservative policy avoids hyperscaler-level storage waste.


Built-in data protection

Native backup integration

VM snapshots are stored directly in S3-compatible object storage. The backup layer is fully independent from the primary block storage layer.

If primary storage is impacted, backups remain intact and accessible.

Locked backups

With S3 object locking (governance or compliance mode), backups cannot be altered or deleted through the whitesky or S3 APIs. Locking is chosen per backup target: no locking, governance or compliance. This protects against:

  • accidental deletion
  • ransomware attacks targeting backup repositories

Rapid recovery

Because backup and storage are integrated, recovery does not depend on external systems. This dramatically reduces recovery time objectives (RTO).


Deployment models

Hyper-converged deployment

Compute and storage run on the same physical servers.

Best suited for:

  • smaller clusters
  • edge locations
  • cost-efficient regional deployments

Benefits:

  • lower hardware footprint
  • simplified operations
  • reduced capital investment

Disaggregated deployment

Dedicated compute nodes and dedicated storage nodes operate independently.

Best suited for:

  • large production environments
  • performance-sensitive workloads
  • independent scaling of compute and storage

This model simplifies failure handling and maintenance at scale.


External storage arrays: bring your own SAN

The built-in distributed storage is the default for new cloud locations. But many organizations moving away from VMware have made significant investments in SAN arrays, have data on them already, and work with strict procurement cycles. For them, replacing the hypervisor should not mean replacing the storage too.

With external storage integration, whitesky runs virtual machines directly on your existing arrays. Your storage stays. VMware goes.

Built on proven Linux technology

There is no proprietary storage layer and no custom distributed filesystem. whitesky combines building blocks that have run in production for decades:

  • LUNs from your array are presented to the servers of a Server-Set, over Fibre Channel or Ethernet. The design works with any SAN that can present LUNs to those servers; compatibility is certified per vendor and SAN type
  • a shared LVM volume group across those servers forms an External Storage Pool
  • SANLock (built into lvm2-lockd) coordinates metadata changes between servers
  • every virtual disk is a logical volume in that pool
External Storage Pool on existing SAN arraysLUNs from any SAN that can present LUNs to the servers of a Server-Set, with compatibility certified per vendor and SAN type, are presented over Fibre Channel or Ethernet and form an External Storage Pool: a shared LVM volume group holding the virtual disks, with SANLock coordinating only metadata changes. All servers of a Server-Set, including a spare server, see all LUNs, so a live migration only moves memory and the virtual disk stays in place.Your SAN arrays, certified per vendor and SAN typevendor Avendor BLUNs over Fibre Channel or EthernetExternal Storage Poolshared LVM volume groupvDiskvDiskvDiskvDiskvDiskvDiskvDiskvDiskSANLock (lvm2-lockd) coordinates metadata only; data I/O goes straight to the arrayServer-Setevery server sees every LUNServerVMVMServerVMVMServerVMVMSpare servertakes the loadduring maintenanceLive migration: only memory moves, the vDisk stays in the pool
External Storage Pool on existing SAN arraysLUNs from any SAN that can present LUNs to the servers of a Server-Set, with compatibility certified per vendor and SAN type, are presented over Fibre Channel or Ethernet and form an External Storage Pool: a shared LVM volume group holding the virtual disks, with SANLock coordinating only metadata changes. All servers of a Server-Set, including a spare server, see all LUNs, so a live migration only moves memory and the virtual disk stays in place.Your certified SAN arraysvendor Avendor BLUNs over FC or EthernetExternal Storage Poolshared LVM volume groupvDiskvDiskvDiskvDiskvDiskvDiskvDiskvDiskSANLock: only metadata is coordinated,data I/O goes straight to the arrayServer-Setevery server sees every LUNServerVMVMServerVMVMServerVMVMSpare servertakes the loadduring maintenanceLive migration: only memory moves,the vDisk stays in the pool
External storage: LUNs from certified SAN arrays form one shared pool for a Server-Set. Live migration moves only memory; the virtual disk never moves.

Only metadata updates need coordination. Data I/O flows directly between the VM host and the storage array: no storage virtualization, no distributed locking and no extra network traffic in the data path. Thin provisioning stays where it belongs, on the array.

Conceptually this is similar to a VMware VMFS datastore, implemented with standard Linux tooling your team already knows.

VMware VMFSwhitesky external storage
ConceptShared datastoreShared volume group
Storage modelBlock storage: SAN LUNsBlock storage: SAN LUNs over Fibre Channel or Ethernet
Live migrationYesYes
CoordinationMetadata lockingSANLock metadata coordination
TechnologyProprietary VMwareStandard Linux LVM
Operational toolingVMware-specificFamiliar Linux stack

Server-Sets: live migration without copying disks

An External Storage Pool is shared by a Server-Set: one or more server blocks whose nodes all see all LUNs. Each server block includes one spare server, so workloads move off a server before maintenance and rolling upgrades run unattended without reducing capacity.

Because source and destination share the same pool, a live migration only transfers VM memory. The virtual disk never moves.

Grow online, mix vendors

  • Online expansion: add a LUN, extend the pool, done. No storage migration, no downtime for servers or running VMs.
  • No vendor lock-in: one pool can hold LUNs from different certified arrays and vendors.
  • Phased array replacement: move from old arrays to new ones incrementally while workloads keep running.
  • Pools of up to 500 virtual disks, by design: smaller pools keep metadata predictable, management operations fast and troubleshooting simple.

Incremental backups with block-level change tracking

The Linux device-mapper module dm-era tracks every 4K block written to a virtual disk. Before each backup run whitesky commits a checkpoint, and only the changed blocks are sent to the backup target. No agent inside the VM, short backup windows, transparent to the workload. Change tracking continues across live migrations, so no changed block is ever missed.

A practical on-ramp for VMware customers

  1. Connect your existing SAN LUNs to the whitesky servers
  2. whitesky registers and validates the External Storage Pool
  3. Migrate VMs from VMware: their disks land on familiar storage
  4. Decommission VMware, while the storage array stays in production
  5. Manage everything through the whitesky portal and API

Customers and operators only see virtual disks. whitesky manages the complexity underneath.

External storage integration was first validated on IBM Storwize arrays, including live migration in the middle of continuous write workloads with verified data integrity. Other arrays are certified per vendor and SAN type. The feature is released and currently rolling out across deployed whitesky cloud locations.


Storage media configurations

Full flash

All layers run on SSD:

  • write buffers
  • distributed transaction logs
  • metadata services
  • erasure-coded storage layer

Delivers:

  • lowest latency
  • consistent performance
  • throughput bounded by network, not disks

Hybrid (flash + HDD)

Flash accelerates hot paths:

  • write buffers
  • metadata
  • cache

HDDs store cold data economically.

Delivers:

  • strong cost-per-TB efficiency
  • flash-like performance for active datasets
  • intelligent background data placement

Both configurations provide identical data safety guarantees.


Optional Security: protection against physical theft

Optional encryption at rest: configured at installation, hosts use encrypted filesystems with keys stored in TPM 2.0 hardware modules. The storage backplane itself is not encrypted.

There are:

  • no centralized key vaults
  • no shared secrets
  • no single point of compromise

Encryption is transparent to workloads and requires no application changes.

Security guarantee, when encryption is configured: Physical theft of disks does not result in data access. Without TPM-secured keys, stolen devices contain only encrypted fragments that cannot be reassembled.


Designed for real-world operations

Maintenance without downtime

Rolling upgrades allow software updates without service interruption. Servers can be replaced transparently while workloads keep running.

Self-healing architecture

Background agents continuously:

  • monitor health
  • repair fragments
  • rebalance capacity
  • verify data integrity

No manual intervention is required.

Operational simplicity

The platform absorbs complexity internally. Operators work with predictable states instead of emergency procedures.

This replaces heroic troubleshooting with calm, repeatable operations.


Technical deep dive (for architects & engineers)

Object storage backend

Core components:

  • OSDs storing object fragments on HDD or SSD
  • Arakoon clusters providing distributed consensus and metadata
  • Namespace managers tracking object locations
  • Stateless proxies for client access
  • Background maintenance agents for repair and rebalancing

Write path:

  • object split into fragments
  • erasure coding applied
  • fragments distributed across fault domains
  • metadata stored in namespace manager

Read path:

  • fragment locations resolved
  • missing fragments reconstructed automatically if needed

Redundancy policies define how many node and disk failures are tolerated.


Virtual block device layer

Virtual disks are exposed via a custom protocol.

Key characteristics:

  • log-structured object aggregation
  • cheap snapshots and clones
  • mapping between logical blocks and objects stored in metadata servers
  • distributed transaction log protects in-flight data

Each virtual disk:

  • is owned by exactly one volume driver
  • can fail over to another driver automatically
  • supports live migration during maintenance

Ownership fencing ensures split-brain conditions cannot corrupt data.


Why customers trust whitesky storage

  • Failure-aware architecture: Built from first principles to isolate, contain, and recover automatically.
  • Sovereignty and control: Transparent operation under your control, aligned with European data sovereignty.
  • Scale without compromise: From edge to datacenter with consistent safety characteristics.

whitesky does not avoid failure — it engineers for it.