Queuack is a pragmatic, single-node job queue that stores jobs in a DuckDB table. Itβs built for dev/test and small-to-medium production workloads where you want durability without the operational overhead of Redis/RabbitMQ/Celery.
Perfect for dev/test environments and small-to-medium production workloads where you want:
- β Persistent queues without Redis/RabbitMQ complexity
- β DAG workflows without Airflow's operational overhead
- β Memory-efficient streaming for processing massive datasets
- β Beautiful visualizations with customizable Mermaid diagrams
- β Zero external dependencies (just DuckDB + stdlib)
- Why Queuack?
- Quick Start
- Documentation & Examples
- Architecture
- Configuration & Tuning
- Security Considerations
- Performance & Scaling
- Testing
- Roadmap
- Contributing
- License
- Acknowledgments
- Support
| Feature | Queuack | Celery + Redis | Airflow | MLflow + Kubeflow |
|---|---|---|---|---|
| Setup | pip install queuack |
Install Redis, configure Celery | Docker compose, PostgreSQL, webserver | K8s cluster, multiple servers |
| DAG Workflows | β Built-in | β Separate tools | β Core feature | β Complex setup |
| Streaming ETL | β O(1) memory | β Load all data | ||
| ML Pipelines | β Native support | β Core feature | ||
| Visualization | β Mermaid (6 themes) | β None | β Web UI (complex) | β Web UI (complex) |
| Local Development | β Single file | β Need K8s | ||
| Memory Footprint | ~50 MB | ~200 MB | ~2 GB | ~4 GB |
Perfect for:
- π¬ ML Engineers - Train models without Kubernetes
- π Data Engineers - Build ETL pipelines without Airflow overhead
- π Startups - Ship fast without infrastructure complexity
- π» Solo Developers - Full workflow engine on your laptop
pip install queuack
# Optional: For Parquet support
pip install queuack[parquet]from queuack import DuckQueue, Worker
# Create queue
queue = DuckQueue("jobs.db") # or ":memory:" for testing
# Enqueue a job
def process_data(x):
return x * 2
job_id = queue.enqueue(process_data, args=(42,))
# Process jobs
worker = Worker(queue, concurrency=4)
worker.run() # Blocks and processes jobsfrom queuack import DAG, DuckQueue
queue = DuckQueue(":memory:")
dag = DAG("etl_pipeline", queue=queue)
# Define tasks
def extract():
return {"records": 1000}
def transform(context):
data = context.upstream("extract")
return {"processed": data["records"] * 2}
def load(context):
result = context.upstream("transform")
print(f"Loaded {result['processed']} records")
return {"status": "success"}
# Build pipeline
dag.add_node(extract, name="extract")
dag.add_node(transform, name="transform", depends_on=["extract"])
dag.add_node(load, name="load", depends_on=["transform"])
# Execute
dag.execute()from queuack import generator_task, StreamReader
# Process 1 million records with ~50MB memory
@generator_task(format="parquet") # or csv, jsonl, pickle
def extract_data():
for i in range(1_000_000):
yield {"id": i, "value": i * 2}
# Returns path to Parquet file
output_path = extract_data()
# Read lazily - one row at a time!
reader = StreamReader(output_path)
for row in reader:
process(row) # Memory stays constantfrom queuack import async_task
import asyncio
# 10-100x speedup for I/O-bound tasks
@async_task
async def fetch_data(urls):
async with aiohttp.ClientSession() as session:
tasks = [session.get(url) for url in urls]
responses = await asyncio.gather(*tasks)
return [await r.json() for r in responses]
# Call synchronously - decorator handles event loop
results = fetch_data(["url1", "url2", ..., "url50"])
# Completes in ~0.1s instead of ~5s (50x faster!)from queuack import DAG, DuckQueue
# Replace Airflow + MLflow + Kubeflow with 50 MB
queue = DuckQueue("ml_pipeline.db")
dag = DAG("model_training", queue=queue)
# Build ML pipeline
dag.add_node(ingest_data, name="ingest")
dag.add_node(validate_data, name="validate", depends_on=["ingest"])
dag.add_node(engineer_features, name="features", depends_on=["validate"])
dag.add_node(train_model, name="train", depends_on=["features"])
dag.add_node(evaluate_model, name="evaluate", depends_on=["train"])
dag.add_node(deploy_model, name="deploy", depends_on=["evaluate"])
# Execute pipeline - all state tracked in SQLite
dag.execute()
# Parallel hyperparameter search
for params in param_grid: # 54 combinations
queue.enqueue(train_model, args=(params,))
# Trains 4x faster with 4 workers, no Ray/Dask needed# Priorities, delays, retries, timeouts
queue.enqueue(
send_email,
args=("user@example.com",),
priority=90, # 0-100 (higher = sooner)
delay_seconds=3600, # Schedule for later
max_attempts=5, # Retry failed jobs
timeout_seconds=300 # Job timeout
)
# Batch operations
job_ids = queue.enqueue_batch([
(task1, (arg1,), {}),
(task2, (arg2,), {}),
])
# Monitor
stats = queue.stats()
# {'pending': 42, 'claimed': 3, 'done': 1250, 'failed': 5}# Fan-out/fan-in pattern
dag.add_node(extract, name="extract")
dag.add_node(transform_a, name="transform_a", depends_on=["extract"])
dag.add_node(transform_b, name="transform_b", depends_on=["extract"])
dag.add_node(load, name="load", depends_on=["transform_a", "transform_b"])
# Conditional execution (ANY mode)
dag.add_node(
validate,
name="validate",
depends_on=["source_a", "source_b"],
dependency_mode="any" # Run when ANY parent completes
)
# Sub-DAGs for reusability
preprocessing_dag = create_preprocessing_dag()
dag.add_subdag(preprocessing_dag, name="preprocess")from queuack import generator_task, StreamReader, StreamWriter
# Write generator β file (O(1) memory)
@generator_task(format="parquet")
def extract():
for i in range(100_000_000): # 100M rows!
yield {"id": i, "data": process(i)}
# Supports 4 formats:
# - JSONL: Human-readable, universal
# - CSV: Excel-compatible
# - Parquet: Analytics, Spark/Pandas
# - Pickle: Complex Python objectsMemory comparison:
- Traditional: Load all β 14 GB RAM β
- Streaming: ~50 MB RAM β
from queuack import MermaidColorScheme
# Pre-built themes
dark = MermaidColorScheme.dark_mode()
professional = MermaidColorScheme.blue_professional()
accessible = MermaidColorScheme.high_contrast()
# Generate diagram
mermaid = dag.export_mermaid(color_scheme=dark)
# Paste into GitHub/GitLab Markdown:
# ```mermaid
# [paste here]
# ```Available themes: default, blue_professional, dark_mode, pastel, high_contrast, grayscale
- β Backpressure control - Automatic throttling at 10k pending jobs
- β Graceful shutdown - SIGTERM/SIGINT handling
- β Dead letter queue - Failed job inspection
- β Claim recovery - Auto-recover stuck jobs
- β Multi-queue workers - Priority-based claiming
- β Concurrent execution - Thread pool workers
- β Transaction safety - ACID guarantees via DuckDB
Our examples follow a progressive learning path:
01_basic/ - Core Concepts
- Simple queue operations
- Priority and delayed jobs
- Batch operations
02_workers/ - Worker Patterns
- Single and concurrent workers
- Multi-queue processing
- Graceful shutdown
03_dag_workflows/ - DAG Patterns
- Linear pipelines
- Fan-out/fan-in
- Conditional execution
- Diamond dependencies
- Sub-DAGs
04_advanced/ - Advanced Techniques
- Custom backpressure
- Monitoring dashboards
- Distributed workers
- Custom Mermaid color schemes
05_real_world/ - Production Use Cases
- ETL pipelines
- Web scraping
- Image processing
- ML training pipelines (see 05_real_world/03_ml/)
- Streaming ETL (1M+ records)
- Multi-format exports (JSONL/CSV/Parquet)
- Async API fetching (50x faster)
06_integration/ - Framework Integration
- Flask, FastAPI, Django
- CLI tools
Run any example:
cd examples/05_real_world
python 01_web_scraper.py- Engine: DuckDB (embedded OLAP database)
- Schema: Single
jobstable with indexes - Locking: File-based for multi-process safety
- Transactions: ACID compliance for atomic operations
- Serialization: Pickle (functions + arguments)
- Concurrency: ThreadPoolExecutor per worker
- Claim Semantics: Visibility timeout with stale recovery
- Retry Logic: Exponential backoff (configurable)
- Graph Engine: NetworkX for topological sorting
- Scheduling: Level-based parallel execution
- Dependencies: ALL (default) or ANY mode
- Status Tracking: Real-time job status monitoring
- Memory Model: O(1) constant memory usage
- Batch Processing: 10k row batches for Parquet
- Format Support: JSONL, CSV, Parquet, Pickle
- Lazy Reading: Generator-based iteration
queue = DuckQueue(
db_path="jobs.db", # or ":memory:"
default_queue="default",
workers_num=4, # Auto-start workers
worker_concurrency=2, # Threads per worker
poll_timeout=1.0, # Claim polling interval
serialization="pickle" # or "json_ref"
)worker = Worker(
queue,
queues=[
("high_priority", 100),
("normal", 50),
("low", 10)
],
concurrency=8, # Thread pool size
worker_id="worker-01" # For distributed setups
)dag = DAG(
name="pipeline",
queue=queue,
max_retries=3,
retry_delay=60,
timeout_per_task=600,
show_progress=True # Progress bar
)# Customize in subclass
class MyQueue(DuckQueue):
@classmethod
def backpressure_warning_threshold(cls):
return 5000 # Warn at 5k pending
@classmethod
def backpressure_block_threshold(cls):
return 50000 # Block at 50k pendingβ οΈ Not safe for untrusted input - Pickle can execute arbitrary codeβ οΈ Not portable across refactors - Function signature changes break old pickles- β Fast and convenient - Works with any Python object
Mitigations:
- Use
serialization="json_ref"mode (functions by reference only) - Validate all job inputs before enqueueing
- Run workers in sandboxed environments
- Keep function signatures stable
- β File-based locking for concurrent workers
- β Automatic stale claim recovery
β οΈ Avoid:memory:with multiple workers (use temp file instead)
- Enqueue: ~5,000 jobs/second (batch mode)
- Claim: ~1,000 claims/second (single worker)
- Execute: Limited by job duration + thread pool
| Jobs/Day | Workers | Concurrency | DB Size |
|---|---|---|---|
| < 10k | 1 | 2-4 | < 100 MB |
| 10k - 100k | 2-4 | 4-8 | 100 MB - 1 GB |
| 100k - 1M | 4-8 | 8-16 | 1-10 GB |
| > 1M | 8+ | 16+ | 10+ GB |
Tips:
- Purge completed jobs regularly (
queue.purge()) - Use multiple queues for different priorities
- Run workers on same host as DB file (avoid network file systems)
- For CPU-bound jobs, use process-based workers
- Monitor with
queue.stats()anddag.get_progress()
# Run all tests
pytest
# With coverage
pytest --cov=queuack --cov-report=html
# Specific test suite
pytest tests/test_dag.py -v
# Fast tests only (skip slow integration tests)
pytest -k "not test_large"- Basic queue with priorities and delays
- DAG workflow engine
- Generator streaming (O(1) memory)
- Multi-format support (CSV, Parquet, JSONL, Pickle)
- Mermaid visualization with themes
- Sub-DAG support
- Async/await support for I/O-heavy tasks
- Job priorities within DAGs
- Terminal UI for monitoring
- Storage backend abstraction (SQLite, PostgreSQL, Redis, S3)
- Scheduled/cron jobs
- Dynamic DAG generation
- Result caching
- Job pause/resume
- Prometheus metrics
We welcome contributions! Please:
- Fork the repository
- Create a feature branch (
git checkout -b feat/amazing-feature) - Write tests for your changes
- Ensure tests pass (
pytest) - Commit your changes (
git commit -m 'feat: Add amazing feature') - Push to the branch (
git push origin feat/amazing-feature) - Open a Pull Request
Development setup:
# Clone repo
git clone https://github.com/brunolnetto/queuack.git
cd queuack
# Install in dev mode
pip install -e ".[dev]"
# Run tests
pytest -vMIT Β© 2025 Bruno Peixoto
- Built with DuckDB - Fast in-process analytical database
- DAG engine inspired by Airflow, simplified for single-node use
- Visualization powered by Mermaid
- Documentation: examples/
- Issues: GitHub Issues
- Discussions: GitHub Discussions
Made with π¦ and β€οΈ by the Queuack team
