Posts

Showing posts from February, 2025

Apache Cassandra Performance Tuning: Data Modeling, Secondary Indexes, and Compaction Strategies

Image
Hello DevOps Engineers, DevSecOps Professionals, SREs, Platform Engineers, and Database Enthusiasts! This week I spent time exploring one of the most popular distributed NoSQL databases— Apache Cassandra . During my hands-on experimentation, I focused on an area that often separates successful Cassandra deployments from struggling ones: performance optimization . Many engineers install Cassandra successfully, but achieving optimal performance requires understanding: Data modeling principles Partition key design Secondary indexes Compaction strategies Read and write optimization In this article, we'll continue from our previous Cassandra installation guide and explore practical techniques to improve query performance and design scalable Cassandra databases. What We'll Learn By the end of this article, you'll understand: How Cassandra data modeling differs from relational databases Why partition keys matter When to use (and avoid) secondary indexes How Cassandra compaction wo...

Cassandra nodetool by examples

Image
To monitor an Apache Cassandra cluster from the command line interface (CLI), you can use the nodetool utility, which is a powerful command-line tool specifically designed for managing and monitoring Cassandra clusters. Here are some key commands and their functionalities: Key nodetool Commands Check Cluster Status : nodetool status This command displays the status of all nodes in the cluster, including whether they are up or down, their load, and other important metrics. Column Family Statistics : nodetool cfstats [keyspace_name] . [table_name] This command provides detailed statistics for a specific table (column family), including read/write counts, disk space used, and more. Thread Pool Statistics : nodetool tpstats This command shows statistics about thread pools used for read, write, and mutation operations, helping to identify potential bottlenecks. Network Statistics : nodetool netstats This command displays information about netwo...

Apache Cassandra 5 installation on Ubuntu

Image
In this post we will have step-by-step process of installation of the Latestt version of Apache Cassandra 5.0.3 (as of Feb 2025 available as latest) on Ubuntu 20. What problem I'm solving with this? There is no direct documentation on the Cassandra to help on installation of latest version that is 5.0.3 on Ubuntu. So I've experimented it on the online Ubuntu terminal(kllercoda) and posting all the steps here. Pre-requisite to install Cassandra 1. Ubuntu Terminal either killercoda or codespace on github works good for this experiment. 2. Cassandra has specific compatibility requirements with different Java versions, which are crucial for ensuring optimal performance and stability. you must have super(root) user access to install cassandra. Step 1: Install Java/JRE Ensure Java installed by checking with command `java -version` and if not existing then install as following: #install jdk17-jre apt install openjdk-17-jre-headless java -version Step 2: Add Ca...

Kafka Message system on Kubernetes

Image
  Setting up the Kubernetes namespace for kafka apiVersion: v1 kind: Namespace metadata: name: "kafka" labels: name: "kafka" k apply -f kafka-ns.yml Now let's create the ZooKeeper container inside the kafka namespace apiVersion: v1 kind: Service metadata: labels: app: zookeeper-service name: zookeeper-service namespace: kafka spec: type: NodePort ports: - name: zookeeper-port port: 2181 nodePort: 30181 targetPort: 2181 selector: app: zookeeper --- apiVersion: apps/v1 kind: Deployment metadata: labels: app: zookeeper name: zookeeper namespace: kafka spec: replicas: 1 selector: matchLabels: app: zookeeper template: metadata: labels: app: zookeeper spec: containers: - image: wurstmeister/zookeeper imagePullPolicy: IfNotPresent name: zookeeper ports: - containerPort: 2181 image1 - kube-kafka1 From th...

Production-Ready Kafka Monitoring with Kafdrop – Complete Setup Guide

Image
Apache Kafka has become the backbone of modern event-driven architectures, powering real-time data pipelines, microservices communication, log aggregation, financial transactions, and IoT workloads. As Kafka clusters grow, monitoring becomes critical to ensure: Topics are healthy Consumer groups are processing messages Lag is under control Partitions are balanced Brokers are available Messages are flowing as expected There are several Kafka monitoring solutions available in the market: There are many several Kafka monitoring tools available, and here I've collected interesting facts about those monitoring tools: Tool Primary Focus Factor House Kpow Enterprise Kafka Operations Datadog Full-stack Observability Logit.io Managed Logging & Monitoring Kafka Lag Exporter Consumer Lag Metrics Confluent Control Center Confluent Platform Monitoring CMAK (Cluster Manager for Apache Kafka) Kafka Administration Kafdrop Lightweight Kafka Web UI Offset Explorer Desktop Kafka Cl...

Kafka installation on Ubuntu Linux

Image
Apache Kafka is one of the most widely adopted distributed event-streaming platforms used by organizations to process, store, and analyze real-time data at scale. Originally developed at LinkedIn and later open-sourced through the Apache Software Foundation, Kafka was designed to handle massive volumes of real-time data generated by millions of users worldwide. In this hands-on tutorial, you will learn how to install Apache Kafka on Ubuntu, configure multiple Kafka brokers on a single host, create topics, publish messages, and consume messages using Kafka command-line tools. What You'll Learn By the end of this tutorial, you will be able to: Understand Kafka architecture and messaging concepts Install Apache Kafka on Ubuntu Linux Configure ZooKeeper for Kafka Create a multi-broker Kafka cluster Create and manage Kafka topics Publish messages using Kafka Producers Read messages using Kafka Consumers Verify Kafka services and network ports What is Apache Kafka? Apache Kafka is a dist...