Historically, inspecting internal Linux kernel behavior or implementing custom packet filtering required writing complex C kernel modules (`.ko`), running severe risks of kernel panic crashes (`Kernel OOPs`) or introducing rootkit security vulnerabilities.

**eBPF (Extended Berkeley Packet Filter)** revolutionizes kernel architecture by allowing developers to run sandboxed bytecode inside the Linux kernel dynamically without modifying kernel source code or loading external kernel modules.

1. The eBPF Virtual Machine Architecture

eBPF introduces a RISC-style virtual machine executing directly inside the Linux kernel with 11 64-bit registers ($R_0$ to $R_{10}$), a 512-byte stack, and zero-overhead JIT (Just-In-Time) compilation to native x86_64 or ARM64 machine instructions.

The Static BPF Verifier

Before any eBPF bytecode program is attached to a kernel hook via the bpf() system call, it must pass the rigorous **BPF Verifier**. The verifier performs static analysis to guarantee kernel safety:

  • Termination Guarantee: Evaluates control flow graphs to prove the program will terminate without infinite loops.
  • Out-of-Bounds Memory Protection: Ensures every pointer dereference is bounds-checked to prevent kernel page faults.
  • Privilege Checks: Restricts privileged helper functions unless loaded by CAP_BPF / CAP_SYS_ADMIN.

2. Tracing Hooks: Kprobes, Uprobes & Tracepoints

Tracing Hook Type Target Execution Subsystem Performance & Use Case
Kprobes / Kretprobes Dynamic kernel symbol entry/exit (e.g. sys_execve). Instruments any arbitrary kernel C function dynamically without recompilation.
Uprobes / Uretprobes Dynamic user-space binary symbols (e.g. OpenSSL SSL_write). Traces user application functions (e.g. capturing unencrypted HTTPS payloads prior to SSL encryption).
Tracepoints Static hardcoded kernel trace instrumentation points. Stable ABI Guarantee: Stable across kernel updates; lower overhead than dynamic kprobes.

3. High-Speed Networking with XDP (eXpress Data Path)

Traditional Linux network processing constructs a complex sk_buff (socket buffer) struct for every incoming packet, passing it through the network stack layers (iptables, routing tables, connection tracking).

**XDP (eXpress Data Path)** executes eBPF programs at the lowest possible layer: *directly inside the Network Interface Card (NIC) driver* before memory allocation for sk_buff occurs:

Possible XDP Return Actions: - XDP_DROP : Instantly drops packet at NIC driver layer (DDoS mitigation at 14M+ packets/sec). - XDP_TX : Bounces packet back out the same NIC network interface. - XDP_REDIRECT: Bypasses host kernel stack and routes packet directly to another virtual NIC or user-space AF_XDP socket. - XDP_PASS : Passes packet up to standard Linux networking stack.

4. eBPF Maps: Sharing Data Between Kernel & User-Space

eBPF programs cannot call standard C library functions or allocate dynamic heap memory. Instead, state is maintained in **eBPF Maps**β€”kernel key-value data structures accessible asynchronously by both kernel eBPF probes and user-space processes (Go, C++, Python):

  • BPF_MAP_TYPE_HASH: Hash table for key-value lookups.
  • BPF_MAP_TYPE_ARRAY: Fast indexed array for metrics counters.
  • BPF_MAP_TYPE_RINGBUF: High-throughput, lockless ring buffer for streaming events to user-space daemons.

5. Working C/Python BCC Tracing Script

Below is a complete, working eBPF tracing script using the BPF Compiler Collection (BCC) in Python. It hooks into the sys_clone kernel function to report every new process created on the host in real time:

#!/usr/bin/env python3 from bcc import BPF # 1. Inline C eBPF Program bpf_text = """ #include struct event_t { u32 pid; char comm[16]; }; BPF_PERF_OUTPUT(events); int kprobe__sys_clone(struct pt_regs *ctx) { struct event_t event = {}; event.pid = bpf_get_current_pid_tgid() >> 32; bpf_get_current_comm(&event.comm, sizeof(event.comm)); events.perf_submit(ctx, &event, sizeof(event)); return 0; } """ # 2. Compile & Attach eBPF Program b = BPF(text=bpf_text) print("Tracing process clone system calls... Press Ctrl+C to exit.\n") print(f"{'PID':<10} {'PROCESS NAME':<20}") # 3. Callback Handler for Perf Ring Buffer def print_event(cpu, data, size): event = b["events"].event(data) print(f"{event.pid:<10} {event.comm.decode('utf-8'):<20}") # 4. Open Ring Buffer Loop b["events"].open_perf_buffer(print_event) while True: try: b.perf_buffer_poll() except KeyboardInterrupt: exit()

6. Modern Enterprise Ecosystem (Cilium & Falco)

  • Cilium: Replaces kube-proxy with eBPF, handling Service Load Balancing, NetworkPolicy enforcement, and mTLS acceleration without iptables latency bottlenecks.
  • Falco: Cloud-native runtime security tool that analyzes eBPF system call streams in real time to detect privilege escalations, unexpected shell spawns, and container breaches.