RisingWave screenshot thumbnail

RisingWave

With over 9,100 GitHub stars and production deployments powering real-time analytics at companies like SHOPLINE where it reduced API latency by 76.7%, RisingWave is the PostgreSQL-compatible streaming database that collapses the traditional Debezium-plus-Kafka-plus-Flink-plus-serving-database stack into a single Rust-powered system. The platform continuously ingests data from PostgreSQL and MySQL via native CDC connectors that eliminate Debezium middleware, consumes Kafka, Redpanda, Pulsar, and Kinesis topics, accepts webhook events from SaaS applications, and batch-loads historical data from S3 and data warehouses. Standard SQL defines sources, materialized views, and sinks — no new DSL, no Java, and no custom API — while the PostgreSQL wire protocol means psql, DBeaver, pgAdmin, Grafana, Metabase, Superset, Tableau, and every PostgreSQL client library works without modification. Materialized views are incrementally maintained as events arrive, delivering point lookups in single-digit milliseconds without recomputing aggregates from scratch. For long-term retention, RisingWave writes to Apache Iceberg tables with a hosted REST catalog and automated table maintenance including compaction, small-file optimization, and snapshot cleanup, with data queryable by Spark, Trino, DuckDB, and DataFusion. The disaggregated compute-storage architecture uses S3-based state management for elastic scaling, instant failure recovery measured in seconds rather than the minutes-to-hours typical of RocksDB-based systems, and cost-efficient storage tiering. An MCP server enables AI agents to query and operate RisingWave directly. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
Apache NiFi screenshot thumbnail

Apache NiFi

Deployed at thousands of enterprises across financial services, healthcare, government, and telecommunications, Apache NiFi is the industry-standard platform for building automated data pipelines through a visual drag-and-drop browser interface that requires zero coding for common integration patterns. The flow-based programming model connects over 300 built-in processors covering relational databases via ExecuteSQL and PutDatabaseRecord, Apache Kafka with PublishKafka and ConsumeKafka, HTTP endpoints through InvokeHTTP and ListenHTTP, cloud storage for AWS S3, Azure Blob, and Google Cloud Storage, SFTP/FTP file transfers, and JSON, XML, CSV, and Avro transformations. Data provenance tracking logs every routing decision, transformation, and delivery for every FlowFile, creating a searchable lineage graph from source to destination with full content replay capability for auditing and debugging. Guaranteed delivery uses configurable backpressure thresholds, prioritized queuing with latency or throughput optimization, and automatic retry with exponential backoff, ensuring no data loss even during downstream outages. The zero-leader clustering architecture distributes processing across nodes with automatic load balancing, while site-to-site protocol enables secure data transfer between NiFi instances across network boundaries. Security includes OpenID Connect and SAML 2.0 single sign-on, role-based access control with fine-grained policies per component, and TLS encryption for all communication. Custom processors can be written in Java and packaged as NAR bundles, or implemented directly in Python through the native scripting framework. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy