76% Faster Replication. Same Infrastructure.
How parallel partitioned reads and write improvements transform enterprise data pipeline performance
CData Sync v26.2 cuts large-scale replication time by 76% on the exact same infrastructure and the exact same workload: a 360 GB, 1-billion-row table into Databricks, no schema changes required.
Reduction in replication time on 360 GB / 1-billion-row datasets
Faster ResultSet metadata retrieval, from 66 ms down to 8 ms
Average write-cycle reduction per file operation on wide tables
The short version, before you download the full report
Enterprise teams that replicate large datasets into cloud warehouses run into the same wall: throughput. As tables grow into the hundreds of gigabytes or billions of rows, replication windows stretch, freshness gaps widen, and cloud compute costs climb. CData Sync v26.x was built to close that gap.
This whitepaper benchmarks two engineering changes shipped in v26.x. Parallel Partitioned Reads splits large source tables into simultaneous multi-threaded reads across available CPU cores. Write Improvements speeds up ResultSet processing and disk-write efficiency on the destination side. Both build on the COPY INTO bulk-loading foundation already used for Snowflake, Databricks, and Redshift.
The headline number: a 360 GB, 1-billion-row SQL Server table that took 11 hours 6 minutes to load into Databricks in v25.1 now completes in 2 hours 41 minutes in v26.2. That's a 76% reduction, on the same infrastructure, with no destination schema changes.
The full report walks through the benchmark methodology, partition-tuning results across four Sync build versions, side-by-side numbers for Snowflake, Databricks, and Redshift, and the exact configuration used to get all three under 3 hours.
Replication speed is a business problem, not just an IT one.
As datasets grow, slow pipelines widen freshness gaps, delay decisions, and burn cloud compute. Three enterprise teams show what's at stake.
Global Pharmaceutical A three-person data team supports 3,500 consumers of SAP financial data. Any delay reaching executive dashboards delays decisions across the entire organization.
Regional Insurance Billions of rows move monthly across 20+ pipelines. A missed replication window pushes the delay directly into next-day business reporting.
Supply Chain Software Replicating 200+ custom ERP fields means per-column processing overhead compounds fast. That's exactly where wide-table write improvements pay off.
Cloud Compute Cost Shorter replication windows cut cost too. They reduce warehouse compute consumption directly in Snowflake, Databricks, and Redshift.
Two engineering advances, one bulk-loading foundation.
Parallel Partitioned Reads
Divides large source tables into simultaneous multi-threaded reads across available CPU cores. On an 8-core Sync machine, 8 parallel partitions cut large-table replication time by 50–75%.
Write Improvements
Faster ResultSet processing and disk-write efficiency on the destination side: an 8% cycle-time reduction per file operation, with the biggest gains on wide tables with hundreds of columns.
COPY INTO Bulk Loading
The cloud-native high-speed ingestion foundation both advances build on, already in place across Snowflake, Databricks, and Redshift.
The full benchmark report
See the numbers behind the 76%
Enter your details and we'll send the full whitepaper: benchmark data, methodology, and the exact configuration that got three major warehouses under 3 hours.