- Grafana Beyla DaemonSet for kernel-level HTTP/gRPC/DNS capture (passive, zero application changes, W3C traceparent header propagation) - flowfish-l7-collector in-cluster bridge: OTLP receiver + buffered pull API - L7 Ingestion Service: K8s service-proxy poll → enrich → RabbitMQ - ClickHouse l7_http_flows / l7_grpc_flows / l7_dns_flows + APM RED MVs - Neo4j L7Workload nodes + SAME_WORKLOAD cross-cluster bridges - New pages: Service Map, Trace Explorer, APM Services List, APM Service Detail - Analysis Wizard now supports L4 / L7 / Both modes with HTTP/gRPC/DNS picks - Integration Hub gains L7 dependency summary + tree-summary integrations - Multi-Cluster Management: dual-agent install (Inspector Gadget L4 + Beyla L7), runtime OpenShift detection so SCCs auto-install with kubectl too - ServiceMap edge → Trace Explorer drill-down with virtual_trace_id correlation - Docs: new L7 architecture diagram, README L7 sections, 3 new screenshots
12 KiB
Flowfish - Technology Choices and Rationale
🎯 Overview
Technologies selected for the Flowfish platform follow principles of high performance, scalability, reliability, and developer productivity.
📊 Technology Stack Summary
| Layer | Technology | Version |
|---|---|---|
| Data Collection (L4) | Inspektor Gadget + eBPF | Latest |
| Data Collection (L7) | Grafana Beyla + eBPF | v3.8+ |
| Backend | Python + FastAPI | 3.11+ / 0.100+ |
| Frontend | React + TypeScript | 18+ / 5+ |
| UI Framework | Ant Design | 5+ |
| Graph Viz | Cytoscape.js | 3.26+ |
| Relational DB | PostgreSQL | 15+ |
| Graph DB | Neo4j | 3.6+ |
| Time-series DB | ClickHouse | 23+ |
| Cache | Redis | 7+ |
| Container | Docker | 20.10+ |
| Orchestration | Kubernetes/OpenShift | 1.27+ / 4.13+ |
🔬 Data Collection: Inspektor Gadget + eBPF
Selection Rationale
Why Inspektor Gadget?
- ✅ Kubernetes Native: Purpose-built for K8s/OpenShift
- ✅ eBPF Powered: Kernel-level data collection, near-zero overhead
- ✅ Zero Application Change: No application changes required
- ✅ DaemonSet Architecture: Easy deployment, automatic scaling
- ✅ Rich Gadget Library: Network, DNS, TCP, process, syscall, file tracking
- ✅ Open Source: MIT license, active community
Alternatives and Why They Were Not Chosen:
| Alternative | Pros | Cons | Why Not Chosen |
|---|---|---|---|
| Service Mesh (Istio/Linkerd) | L7 metrics, mTLS | Sidecar injection required, high overhead | Requires application changes |
| APM Tools (Datadog, New Relic) | Rich UI, easy setup | Paid, vendor lock-in | Costly, external dependency |
| Custom eBPF Programs | Full control | Development complexity | High development cost |
| Network Sniffer (tcpdump) | Simple | High performance impact | Scalability issues |
Technical Details
eBPF (Extended Berkeley Packet Filter):
- Linux kernel 4.4+ support
- In-kernel execution (no user-space round trips)
- Verifiable bytecode (safe execution)
- CO-RE (Compile Once Run Everywhere)
- Minimal CPU/memory overhead (<1–2%)
Inspektor Gadgets:
trace_network: TCP/UDP connection trackingtrace_dns: DNS query/response loggingtrace_tcp: TCP lifecycle eventstrace_exec: Process execution trackingtrace_open: File access monitoringtrace_bind: Port binding detection
🔬 L7 Data Collection: Grafana Beyla + eBPF
Selection Rationale
Why Grafana Beyla?
- ✅ eBPF-Based: Same kernel-level approach as Inspektor Gadget
- ✅ L7 Protocol Support: HTTP, gRPC, DNS request/response capture
- ✅ Zero Application Change: No sidecars, no code instrumentation
- ✅ DaemonSet Architecture: Consistent with existing Inspektor Gadget deployment
- ✅ Multi-Arch: AMD64 + ARM64 support (~50MB image)
- ✅ OpenTelemetry Native: Exports OTLP traces and metrics
- ✅ Apache 2.0 License: Fully open source
Alternatives and Why They Were Not Chosen:
| Alternative | Pros | Cons | Why Not Chosen |
|---|---|---|---|
| Kubeshark | Rich L7 capture, UI | License changed to paid, limited multi-cluster | Commercial licensing |
| Cilium Hubble | eBPF, L7 visibility | Requires Cilium CNI | Not CNI-agnostic |
| Pixie | Rich L7 visibility | Cloud-centric, limited self-hosted | External dependency |
| Service Mesh | mTLS, L7 metrics | Sidecar injection, high overhead | Requires application changes |
Beyla Architecture:
- Deployed as DaemonSet on each cluster node
- Instruments Go, Python, Java, Node.js, .NET, Rust applications automatically
- Captures HTTP method/path/status, gRPC service/method, DNS queries
- Exports OpenTelemetry spans to in-cluster flowfish-l7-collector
- flowfish-l7-collector bridges push model to Flowfish pull model via K8s API Service Proxy
🚀 Backend: Python + FastAPI
Selection Rationale
Python:
- ✅ Ecosystem: Rich library support (data processing, ML/AI)
- ✅ LLM Integration: Libraries like OpenAI, LangChain are native to Python
- ✅ Async Support: Modern async programming with asyncio
- ✅ Developer Productivity: Fast development, readable syntax
FastAPI:
- ✅ High Performance: Starlette + Pydantic, Go/Node.js–level speed
- ✅ Automatic OpenAPI: Swagger UI generated automatically
- ✅ Type Safety: Compile-time type checking with Pydantic
- ✅ Async Native: Native async/await support
- ✅ Dependency Injection: Clean, testable code
- ✅ WebSocket Support: Real-time communication
Alternatives:
| Alternative | Why Not Chosen |
|---|---|
| Django | Monolithic, overhead for non-REST needs |
| Flask | Sync-only, lacks modern features |
| Go (Gin/Echo) | Weaker Python ecosystem and LLM integration |
| Node.js (Express) | Callback hell, weak type safety |
⚛️ Frontend: React + TypeScript + Ant Design
Selection Rationale
React 18:
- ✅ Industry Standard: Large community, abundant resources
- ✅ Component-Based: Reusable, maintainable components
- ✅ Hooks: Modern state management
- ✅ Virtual DOM: Efficient rendering
- ✅ Server Components: Future-proof (RSC)
TypeScript:
- ✅ Type Safety: Catch errors at compile time
- ✅ Better IntelliSense: Excellent IDE support
- ✅ Refactoring: Safe rename, move operations
- ✅ Documentation: Types = self-documenting code
Ant Design (antd):
- ✅ Enterprise-Grade: Used by Fortune 500 companies
- ✅ Comprehensive: 60+ high-quality components
- ✅ Consistent: Unified design language
- ✅ Customizable: Theme support, CSS-in-JS
- ✅ Accessible: WCAG 2.0 AA compliant
- ✅ I18n: Multi-language support built-in
Alternatives:
| Alternative | Why Not Chosen |
|---|---|
| Vue.js | Smaller ecosystem, less enterprise adoption |
| Angular | Steep learning curve, verbose |
| Material-UI | Ant Design is more enterprise-focused |
| Chakra UI | Younger, less battle-tested |
🎨 Graph Visualization: Cytoscape.js
Selection Rationale
Cytoscape.js:
- ✅ Purpose-Built: Designed specifically for graph visualization
- ✅ Performance: Handles 1000+ nodes/edges
- ✅ Extensible: Plugin ecosystem
- ✅ Layout Algorithms: Hierarchical, force-directed, circular, grid
- ✅ Styling: CSS-like styling system
- ✅ Events: Rich interaction events
- ✅ Export: PNG, JPG, JSON export
Alternatives:
| Alternative | Pros | Cons |
|---|---|---|
| D3.js | Very flexible, powerful | Steep learning curve, verbose |
| Vis.js | Easy to use | Performance issues (>500 nodes) |
| Sigma.js | Fast rendering | Limited feature set |
| React Flow | React-native | Missing graph algorithms |
🗄️ Databases
PostgreSQL 15+ (Relational Data)
Selection Rationale:
- ✅ ACID Compliance: Reliable transactions
- ✅ JSONB Support: JSON storage for flexible schema
- ✅ Full-Text Search: Built-in search capabilities
- ✅ Extensions: PostGIS, pg_trgm, btree_gin
- ✅ Replication: Streaming replication, logical replication
- ✅ Partitioning: Table partitioning for large datasets
- ✅ Mature: 30+ years, production-proven
Use Cases:
- User accounts, roles, permissions
- Cluster and namespace metadata
- Analysis configurations
- Anomaly and change records
- Audit logs
Why Alternatives Were Not Chosen:
- MySQL: Weak JSONB support, complex replication
- MongoDB: Weak ACID guarantees, not ideal for relational data
Neo4j 3.6+ (Graph Database)
Selection Rationale:
- ✅ Distributed: Native distributed architecture
- ✅ Scale: Support for trillions of vertices/edges
- ✅ Performance: Sub-millisecond graph traversal
- ✅ GQL (nGQL): SQL-like graph query language
- ✅ Open Source: Apache 2.0 license
- ✅ Kubernetes-Friendly: Helm charts, operators
- ✅ Consistency: Strong consistency via Raft
Use Cases:
- Workload dependencies (Pod → Deployment → Service)
- Communication edges (COMMUNICATES_WITH)
- Dependency chains (DEPENDS_ON)
- Graph traversal queries (upstream/downstream)
Alternatives:
| Alternative | Why Not Chosen |
|---|---|
| Neo4j | Paid (enterprise), Cypher proprietary |
| JanusGraph | Lower performance than Neo4j |
| Amazon Neptune | Vendor lock-in, cloud-only |
| ArangoDB | Multi-model complexity |
ClickHouse 23+ (Time-Series/OLAP)
Selection Rationale:
- ✅ Columnar Storage: High compression (10–100x)
- ✅ Fast Queries: Billions of rows, sub-second queries
- ✅ Aggregations: Pre-aggregation via materialized views
- ✅ TTL Support: Automatic data cleanup
- ✅ Partitioning: Date/time based partitioning
- ✅ Replication: Built-in replication
- ✅ SQL: Standard SQL dialect
Use Cases:
- Network flow events (raw eBPF data)
- DNS queries, TCP connections
- HTTP requests, metrics
- Process events, syscall traces
- Aggregated request metrics
Alternatives:
| Alternative | Why Not Chosen |
|---|---|
| TimescaleDB | PostgreSQL extension, slower |
| InfluxDB | Non-SQL, limited query capabilities |
| Elasticsearch | Resource-heavy, complex operations |
| Prometheus | Short retention, not for raw events |
Redis 7+ (Cache & Real-time)
Selection Rationale:
- ✅ In-Memory: Microsecond latency
- ✅ Pub/Sub: Real-time event streaming
- ✅ Data Structures: Lists, sets, sorted sets, hashes
- ✅ TTL: Automatic expiration
- ✅ Persistence: RDB + AOF
- ✅ Clustering: Native clustering support
- ✅ Sentinel: Automatic failover
Use Cases:
- Session storage (JWT tokens)
- Real-time metrics cache
- Rate limiting counters
- Pub/Sub for WebSocket updates
- Distributed locks
🐳 Container & Orchestration
Docker
Selection Rationale:
- ✅ Industry Standard: De facto containerization platform
- ✅ Image Registry: Docker Hub, private registries
- ✅ Multi-Stage Builds: Optimized images
- ✅ BuildKit: Fast, efficient builds
Kubernetes / OpenShift
Selection Rationale:
- ✅ Cloud-Native Standard: Industry standard orchestration
- ✅ Auto-Scaling: HPA, VPA
- ✅ Self-Healing: Automatic restarts, health checks
- ✅ Service Discovery: Built-in DNS
- ✅ Storage: PersistentVolumes, StorageClasses
- ✅ Security: RBAC, NetworkPolicies, PodSecurityPolicies
- ✅ OpenShift: Enterprise features, operators, built-in monitoring
🔐 Authentication
JWT (JSON Web Tokens)
Selection Rationale:
- ✅ Stateless: No server-side session required
- ✅ Scalable: Horizontal scaling friendly
- ✅ Cross-Domain: CORS-friendly
- ✅ Standard: RFC 7519
- ✅ Libraries: Mature libraries for every language
OAuth 2.0 / OpenID Connect
Selection Rationale:
- ✅ SSO: Single Sign-On support
- ✅ Enterprise: Azure AD, Okta, Keycloak integration
- ✅ Delegation: Secure delegation of access
- ✅ Standard: Industry standard protocol
📊 Monitoring & Observability
Recommended Stack (Optional)
| Component | Technology | Purpose |
|---|---|---|
| Metrics | Prometheus | Time-series metrics |
| Logs | Loki / ELK | Centralized logging |
| Tracing | Jaeger / Tempo | Distributed tracing |
| Dashboards | Grafana | Visualization |
| Alerting | Alertmanager | Alert management |
🎯 Conclusion
The Flowfish technology stack follows modern cloud-native application development best practices:
✅ Performance: eBPF, FastAPI, ClickHouse, Redis
✅ Scalability: Kubernetes, distributed databases
✅ Reliability: PostgreSQL ACID, replication
✅ Developer Experience: Python, TypeScript, React
✅ Maintainability: Type safety, test frameworks
✅ Open Source: No vendor lock-in, community support
Version: 1.0.0
Last Updated: January 2025