A high-performance C++20 coroutine library designed for batch task processing and high-throughput scenarios.
English | 中文
- High Performance: Optimized for high-throughput scenarios with Profile-Guided Optimization (PGO)
- Lock-free Architecture: Lock-free queues, a single coroutine scheduler by default, plus a thread pool for blocking/CPU work
- C++20 Coroutines: Modern coroutine-based task scheduling
- Batch Processing: Designed for concurrent task execution
- Advanced Concurrency: WhenAny, WhenAll operations with optimized scheduling
- Channel Communication: Thread-safe async channels for producer-consumer patterns
- Memory Pool: Custom memory allocation inspired by Redis/Nginx design
- PGO Optimization: Performance improvements through profile-guided compilation
- Deterministic Real-Time Execution: Single-thread-affine
RtExecutorfor robotics / control / embedded — periodic ticks, CPU pinning, and a measured jitter lamp (flowcoro::rt). See the rt control-loop demo; the DDS pipeline demo usesTask<>, notRtExecutor. - Bounded Lock-Free MPMC Channel:
BoundedChannel<T>(Vyukov ring — no allocation, no SMR, immediate fail when full/empty) - CUDA stream awaiters:
PinnedHostBuffer/DeviceBuffer/co_await streamviacuLaunchHostFunc(used by flowtrain’s async 1F1B). Learn: docs/LEARN_CUDA.md - CPU Affinity:
cpu_affinity.hunifies pinning and physical-core enumeration (dedup bythread_siblings_list, so SMT siblings never go to two workers);lockfree::ThreadPooltakes an affinity list directly - Python Bindings (optional):
-DFLOWCORO_BUILD_PYTHON=ONbuildsflowcoro_py.so—CoroutineThreadPool+when_all/wait_any+Channel; see Python binding docs
Measured numbers live in Performance Data Reference. The figures below are from a 2026-09-15 Release run of professional_flowcoro_benchmark on the author's machine (Intel Core i7-14650HX 6C/12T, Ubuntu 24.04.4 LTS WSL2, g++ 13.3.0, FlowCoro 4.0.0, thread_count=12). They are not a claim of production-proven scale.
| Benchmark | Mean | Throughput |
|---|---|---|
| Simple Computation | 14 ns | 69.0M ops/s |
| Coroutine Create & Execute | 132 ns | 7.60M ops/s |
| WhenAny (2 tasks) | 630 ns | 1.59M ops/s |
| LockFree Queue (enq+deq) | 132 ns | 7.56M ops/s |
| Memory Allocation (1KB) | 24 ns | 41.3M ops/s |
Some bench labels (HTTP, Echo) are CPU-side simulations, not real sockets. This run did not refresh Go numbers; a same-machine Rust microbench exists but is not the same methodology — see PERFORMANCE_DATA.md rather than “N× faster than Go/Rust” claims.
Best suited for: coroutine scheduling, batch / high-throughput work, plus deterministic realtime via flowcoro::rt. Sibling LLM testbeds: flowserve (serving schedule), flowtrain (DP/TP/PP mechanism). This repo is not an LLM trainer.
- C++20 compiler (GCC 11+, Clang 12+)
- CMake 3.16+
- Linux/macOS/Windows
git clone https://github.com/caixuf/flowcoro.git
cd flowcoro
mkdir build && cd build
cmake .. -DCMAKE_BUILD_TYPE=Release
make -j$(nproc)Note: PGO optimization is automatically enabled when profile data is available in pgo_profiles/ directory.
#include <flowcoro.hpp>
using namespace flowcoro;
// Simple coroutine task
Task<int> compute(int value) {
// Simulate work
co_await sleep_for(std::chrono::milliseconds(10));
co_return value * 2;
}
// Concurrent execution
Task<void> example() {
// Tasks start executing synchronously, then are scheduled when they suspend
auto task1 = compute(10);
auto task2 = compute(20);
auto task3 = compute(30);
// Wait for results
auto result1 = co_await task1;
auto result2 = co_await task2;
auto result3 = co_await task3;
std::cout << "Results: " << result1 << ", " << result2 << ", " << result3 << std::endl;
co_return;
}
// Producer-Consumer with Channel
Task<void> channel_example() {
auto channel = make_channel<int>(10); // Buffer size: 10
// Producer
auto producer = [channel]() -> Task<void> {
for (int i = 0; i < 5; ++i) {
co_await channel->send(i);
std::cout << "Produced: " << i << std::endl;
}
channel->close();
};
// Consumer
auto consumer = [channel]() -> Task<void> {
while (true) {
auto value = co_await channel->recv();
if (!value.has_value()) break; // Channel closed
std::cout << "Consumed: " << value.value() << std::endl;
}
};
auto prod_task = producer();
auto cons_task = consumer();
co_await prod_task;
co_await cons_task;
co_return;
}
int main() {
sync_wait(example());
sync_wait(channel_example());
return 0;
}pip install pybind11
cmake -B build-py -DCMAKE_BUILD_TYPE=Release -DFLOWCORO_BUILD_PYTHON=ON
cmake --build build-py --parallel $(nproc) --target flowcoro_py
export PYTHONPATH=$PWD/build-py/pythonimport flowcoro_py as fc
with fc.CoroutineThreadPool(threads=8, pin_to_cores=True) as pool:
futures = [pool.submit(run_case, path) for path in cases]
results = fc.when_all(futures) # order-preserving; raises the lowest-index failure
channel = fc.Channel(capacity=10240) # bounded lock-free MPMC, bytes payload
channel.try_push(b"pdu") # -> bool, non-blocking
msg = channel.try_pop() # -> Optional[bytes], non-blockingRead docs/PYTHON_BINDING.md first. submit runs Python
callables on the C++ pool, so only work that releases the GIL (subprocess / IO) gets real
parallelism — measured against concurrent.futures.ThreadPoolExecutor it is a wash on the
shapes that matter, and slower on micro-tasks. The binding is worth it for precise core
pinning and an in-process bounded lock-free hand-off, not as a magic speedup.
FlowCoro uses a three-layer scheduling architecture:
Task Creation → Coroutine Manager → Coroutine Pool → Thread Pool
↓ ↓ ↓ ↓
suspend_never Load Balancing Lock-free Queue Execution
- suspend_never: Tasks start executing synchronously on creation, then suspend at first co_await point
- Smart Load Balancing: Automatic scheduler selection after suspension
- Lock-free Queues: High-performance task distribution
- Continuation Mechanism: Zero-copy task chaining
Alongside the high-throughput scheduler, FlowCoro ships a single-thread deterministic real-time model (RtExecutor) for latency-critical control loops (robotics / autonomous driving / embedded):
- Single-thread affinity: every
resume/destroyhappens on the executor thread; cross-thread events onlypost_ready(h)the handle back — never inlineresume. This is what FlowEngine depends on to avoid data races. - Periodic tick:
run()is a non-blocking tick (timers → ready → resume/destroy).rt::sleep_untilaligns to an absolute deadline;rt::yield()defers to the next tick. - CPU pinning:
Config{.pin_cpu = N}binds the run() thread on first tick (apply_affinity(), Linux). - Local hot path: internal re-post uses a reserved thread-local vector. Cross-thread
post_readyallocates a queue node;run_blockingmaysleep_untilwhen idle. That is not “zero syscall everywhere”. - Two-phase shutdown:
request_stop()→ cooperative stop at a cycle boundary → framesco_returnand are destroyed on the executor thread.
#include <flowcoro/rt_executor.h>
flowcoro::rt::RtExecutor ex{{ .pin_cpu = 2 }};
ex.spawn(periodic_control_task(), "control");
ex.run_blocking(); // periodic ticks until all tasks finish
// or: while (!ex.is_finished()) ex.run();See the contract in API Reference §9, jitter lamp tests/test_rt_latency.cpp, and the rt control-loop demo. The DDS pipeline demo is Task+Channel middleware, not RtExecutor.
Ideal for:
- Web API servers (high request throughput)
- Batch data processing
- High-frequency trading systems
- Microservice gateways
- Producer-consumer pipelines
- Multi-stage data processing workflows
With the Real-Time layer (flowcoro::rt):
- Robotics / autonomous driving control loops (periodic, latency-critical)
- AMR / drone stabilization and navigation
- Embedded deterministic scheduling
- Any embedded/control workload needing single-thread affinity + CPU pinning
Try the rt control-loop demo:
cd build && cmake .. && cmake --build . --target rt_control_loop_demo
./examples/autonomous_driving/rt_control_loop_demo 2Enhanced with Channels:
- Inter-coroutine communication
- Buffered message passing
- Async data pipelines
- Coordinated multi-producer/multi-consumer systems
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DFLOWCORO_BUILD_BENCHMARKS=ON
cmake --build build --target professional_flowcoro_benchmark -j$(nproc)
./build/benchmarks/professional_flowcoro_benchmark
# Optional scale checks (different from the professional bench)
./build/examples/hello_world 10000 # 10K tasks
./build/examples/hello_world 100000 # 100K tasks- Introduction (中文)
- Quick Start Guide
- API Reference
- Architecture Design
- Performance Data
- PGO Optimization Guide
- Python Bindings
- CUDA awaiters & BoundedChannel (for flowtrain)
MIT License - see LICENSE file for details.
Contributions are welcome! Please read our contributing guidelines and submit pull requests.