In an ETL setup, real-time streaming integration keeps a continuous stream of data moving from source systems as each event occurs. Unlike batch ETL or scheduled micro-batch jobs, it doesn’t wait to move data in chunks. Most data teams who’ve gone past a pilot or even a prototype already know which pattern they want, whether it’s CDC, event streaming, or change polling.
While these patterns are well documented, building the connector to the actual source feeding that pattern isn’t easy in production, with millions of rows moving every day. It gets even harder when that source is a proprietary database, a legacy ERP, or a custom REST API with no off-the-shelf support.
To solve that problem, this guide lays out the core streaming patterns, why custom sources break them, and how a standards-based connector layer makes those patterns work against any source.
What is real-time streaming integration?
Quick definition: Real-time streaming integration
A data integration pattern for the continuous capture, processing, and movement of data from source systems to destinations in ‘real-time’ as each event or data change occurs.
It’s called real time because stream processing handles data as soon as it arrives, with latency ranging from sub-second to a few seconds.
Event streaming, where producers publish events and consumers read them as they happen, sits underneath most streaming integration layers. In simple terms, it’s a running log of events, such as an order placed or a row updated. Platforms like Apache Kafka let systems publish, store, and process these streams of events.
Real-time integration is great for operational workloads that can’t wait for the next batch run. Data and engineering teams adopt the streaming pattern when the value of the data drops within seconds or minutes, such as in:
Fraud detection on payment and transaction events
Real-time ad optimization and cybersecurity threat detection
Operational monitoring and live inventory tracking
IoT and sensor monitoring
Each of these use cases depends on immediate insight rather than a nightly refresh. For a broader look at the real-time streaming pattern, see our guide on real-time data.
Streaming integration vs. batch ETL: key differences
Quick definition: Batch ETL
A data integration approach that extracts, transforms, and loads accumulated data in scheduled chunks or ‘batches’ at defined intervals, rather than continuously.
Batch ETL collects data over time and processes it in scheduled runs, while streaming integration processes each record the moment it arrives.
Choosing between the two comes down to these key dimensions:
Dimension | Batch ETL | Streaming integration |
Latency | High; data waits for the next scheduled run | Low; sub-second to a few seconds |
Processing unit | Accumulated dataset per run | Individual events or changes |
Infrastructure | Simpler; scheduled compute that shuts down after each run | More complex; state management, ordering, and fault tolerance |
Cost to operate | Predictable; compute runs only during jobs | Continuous; always-on infrastructure |
Error recovery | Fix the bug and rerun the batch | Checkpointing, replay, and careful state recovery |
Best for | Reporting, historical analysis, large backfills | Fraud detection, live dashboards, CDC replication |
While batch is simpler and cheaper to run, streaming closes the latency gap but costs more to operate and recover. This isn’t a winner-takes-all choice, and most production environments run both, depending on the workload. Our post on real-time processing covers the processing side in more depth.
Running both doesn’t have to mean two tools. CData Sync can schedule replication jobs at any interval for batch loads, or run them in real time with CDC.
Core real-time streaming integration patterns
Most real-time streaming setups come down to four distinct patterns, and each one solves a different part of the problem. They aren’t rigid, and many pipelines combine two or more, such as CDC feeding an event stream.
1. Event streaming
Event streaming is a continuous flow of events from producers to consumers through a durable log. In Kafka, producers write events to topics, and consumers read from those topics on their own schedule.
Producers and consumers never talk to each other directly, so each side scales and recovers independently.
2. Log-based change data capture (CDC)
Quick definition: Change data capture (CDC)
A technique that identifies and extracts only the data that changed in a source system, including inserts, updates, and deletes, so downstream systems stay in sync without having to replicate entire datasets again.
Log-based CDC reads a database’s native transaction log to capture every insert, update, and delete as it happens. Common examples are the PostgreSQL write-ahead log (WAL), the MySQL binlog, and the Oracle redo log.
The database already writes these logs for recovery, so reading them is non-intrusive. The CDC process tails the log instead of querying production tables and gets a complete, ordered record of changes.
3. Change polling (query-based CDC)
Change polling runs a query on a schedule to find rows that changed since the last run, using a timestamp or version column. It’s simple to set up and works well for smaller or legacy systems.
Polling has two known limits. It puts a constant, repeated load on the source, and it usually misses deletes because a deleted row is no longer there to query.
4. Hub-and-spoke integration
Hub-and-spoke integration connects each system once to a central hub instead of wiring every system to every other system. Connecting N systems directly needs N(N-1)/2 links, while a hub needs only N.
At 10 systems, that’s 45 direct links compared to 10 hub connections. The hub also gives you one place for monitoring, logging, and governance of integration flows. In streaming setups, an event broker like Kafka often plays the hub role.
Pattern comparison
Pattern | How it works | Latency | Best for | Main tradeoff |
Event streaming | Producers publish to a durable log; consumers subscribe | Low | Decoupled flows with many consumers | Running and operating the log infrastructure |
Log-based CDC | Reads the transaction log | Lowest | Real-time database replication | Needs log access; not all sources expose it |
Change polling | Queries for rows changed since the last run | Bound by poll interval | Sources without log access, including APIs | Adds source load; often misses deletes |
Hub-and-spoke | Each system connects once to a central hub | Depends on the hub | Many systems exchanging data | The hub becomes a critical dependency |
Why custom and proprietary sources are uniquely challenging
Connectors for proprietary sources bring in additional challenges while implementing streaming patterns. While PostgreSQL and MySQL have mature CDC support, not every database exposes its transaction log the same way, and some proprietary databases need vendor-specific tools or don’t support log-based CDC at all.
Without log access, teams fall back to polling, which puts steady, repeated load on the source and typically misses deletes, so the pipeline keeps running while the destination slowly drifts from the source.
Custom REST and SaaS APIs bring even more work:
Pagination: Cursor, offset, and link-header schemes differ by vendor.
Authentication: OAuth flows, token refresh, and API keys need handling in every connector.
Rate limits: Polling too often gets you throttled; polling too rarely raises latency.
Schema drift: New fields and changed types break hard-coded mappings.
Each of these ends up inside a hand-coded connector needing updates whenever the vendor changes its API, and the bigger issue for integration architects is coupling, since custom trigger logic ties application code to data infrastructure and makes maintenance harder as schemas change, as the Salesforce to Apache Kafka pattern shows for even a common CRM source.
A reusable, standards-based connector layer keeps that logic out of the application. CData Sync compares source and destination schemas on every run and automatically alters the destination for new columns or wider types.
How ETL tools handle custom sources with standards-based connectors
Standards-based connectors expose proprietary sources through common interfaces like ODBC, JDBC, and REST, so a pipeline queries any source through the same code path. Open Database Connectivity (ODBC) is a widely accepted API built so a single application can reach different systems with the same source code. Java Database Connectivity (JDBC) does the same for any Java-based tool, such as Kafka Connect, Flink, or Spark.
The driver is the translation layer. It gives the application a standard set of functions and handles the source-specific work underneath, which is how a proprietary source can look like a standard SQL source to a streaming tool.
Standard | Primary environment | What it abstracts | Example use in streaming |
ODBC | C/C++, BI, and ETL tools | Source APIs and dialects behind one SQL interface | Polling a legacy ERP from a scheduled ingestion job |
JDBC | Java/JVM tools | Relational and non-relational sources behind one API | Feeding a Kafka Connect or Flink job from a custom source |
REST (real-time web services) | Any HTTP client | Endpoints, auth, and response formats behind a consistent API | Exposing a custom source to serverless functions or microservices |
CData Sync builds on the same idea, with hundreds of pre-built connectors that generate dynamic schemas and table views even for sources that don’t expose them natively.
Build vs. buy: choosing serverless and hosted integration for custom streaming pipelines
Once the pattern and source are clear, the build-vs-buy decision for custom streaming connectors comes down to a few factors: how many sources you need to reach, who maintains the connectors, and where the pipeline runs. The matrix below maps these out.
Quick definition: Serverless integration
A data integration model where the cloud provider fully manages the underlying servers, with automatic scaling, built-in high availability, and pay-for-use billing, so teams focus on logic instead of infrastructure.
Serverless fits streaming well on the event side, with AWS describing its serverless technologies as having automatic scaling, built-in high availability, and pay-for-use billing, and services like AWS Lambda running code in response to events and reading directly from Kafka and Kinesis streams. It works best for short, event-triggered tasks, though containers or VMs may be the better choice for applications with many continuous, long-running processes, which matters when you size an always-on stream.
Hosted integration offers a middle path, where managed Kafka services like IBM Event Streams handle durability and high availability so teams build applications instead of running brokers.
Approach | You manage | Best when | Tradeoff |
Build: hand-coded connectors | Connector code, auth, pagination, schema changes, and pipeline logic | One or two sources with unusual requirements | Full control, but maintenance grows with every source and API change |
Build: serverless functions | Function logic and connector code | Short, event-triggered workloads | Less suited to long-running streams; connector code is still yours |
Buy: hosted streaming service | Pipeline logic and connectors | Teams that want Kafka without running brokers | Solves infrastructure, but custom source connectors are still your job |
Buy: ETL tool with pre-built connectors | Pipeline logic only | Many custom or long-tail sources to reach fast | Less low-level control than fully hand-coded |
Hosted and serverless platforms solve the infrastructure problem. Pre-built connectors solve the source problem, and the two work together.
CData Sync covers the “buy” side for the connector layer and runs wherever your pipeline does: on-premises, in AWS, Azure, or GCP, or as CData-hosted Sync Cloud. Orchestration and stream processing stay on your existing platform. For a wider view, see our roundup of data ingestion tools.
Build reliable real-time pipelines from custom sources with CData Sync
Reliable real-time pipelines from custom sources come from pairing a proven streaming pattern with a reusable connector layer. One standard interface lets a single pipeline reach many proprietary sources with the same code.
CData Sync supports log-based CDC for MySQL, MariaDB, Oracle, PostgreSQL, and SQL Server, plus change tracking for Microsoft Dynamics 365. For hundreds of other sources, it uses incremental check columns to pull only new or changed records.
You can start a free trial of CData Sync to test it against your own source. For more on how these pieces fit together, read our overview of data integration.
Frequently asked questions
What is real-time streaming integration, and how is it different from batch ETL?
Real-time streaming integration captures and delivers data continuously as each event occurs. Batch ETL collects data over an interval and processes it in scheduled runs or ‘batches.’ Streaming gives lower latency, while batch is simpler and cheaper to operate.
What are the most common real-time streaming integration patterns for custom data sources?
The four core patterns are event streaming, log-based CDC, change polling, and hub-and-spoke integration. For custom sources without log access, change polling through a standards-based driver is often the practical starting point.
How does change data capture (CDC) work, and when should you use it?
Log-based CDC reads the source database’s transaction log to capture every insert, update, and delete in commit order. Use it when the source exposes its log and you need low latency, delete capture, and minimal load on production tables.
What technologies are commonly used to build a real-time streaming integration pipeline?
A typical pipeline combines an event streaming platform like Apache Kafka, a CDC or ingestion tool, a stream processor like Apache Flink, and connectors to sources and destinations. ODBC, JDBC, and real-time web services (REST) handle source connectivity.
What’s the difference between event streaming and stream processing?
Event streaming moves and stores events between producers and consumers. Stream processing transforms, filters, joins, or aggregates those events as they flow. Kafka handles event streaming, while tools like Kafka Streams or Flink handle processing.
When should you build a custom source connector versus using a hosted or serverless integration platform?
Build when you have one or two sources with unusual requirements and a team to maintain them. When you need to reach many custom sources, use hosted or serverless platforms for infrastructure and pre-built connectors for the source layer.
What latency should you expect from real-time streaming with a custom source?
Typical CDC latency ranges from sub-second to a few seconds, depending on network, log parsing, and serialization overhead. Polling pipelines are bound by the poll interval and the source’s API rate limits.