Rafael Reis

Rafael
Reis.

Platform & Infrastructure Engineer · Engineering Manager
LocationSão José dos Campos, BR
FocusSRE, DevOps, AI infra
StatusAutomatic failover, nobody woken up
LanguagesPT · EN · ES
00

Abstract

I build and run multi-tenant SaaS infrastructure. The entire platform team, from the rack to the runtime.

Hands-on platform engineer (SRE · DevOps), engineering lead, and co-founder of a lean multi-tenant SaaS. I design, build and keep production alive, and I have led the teams that do it. Plus self-hosted AI on my own metal.

60+production tenants 50+migrated, zero downtime 40+self-hosted services 12+years of homelab
01

Capabilities

Platform & Infrastructure

Linux, Proxmox/KVM, Docker, GitOps, identity (OAuth 2.0/OIDC, RBAC), webhooks and APIs, plan/apply tenant onboarding. I build the thing that everything else runs on.

Reliability (SRE)

HA & failover: Patroni, replica sets, Sentinel. Observability, backups & DR, incident response with real postmortems.

AI Infrastructure

Self-hosted LLMs on GPU (llama.cpp), RAG and embeddings, LLMOps. AI that runs on my own metal.

Operations & engineering leadership

Teams of 10+ engineers, a 40+ person delivery network, marketplace P&L at Rappi, a KPI turnaround at Contabilizei. I get the business the systems serve, and the people who run them.

02

Selected work

architecture, decisions, the occasional 3am incident

Migrating 50+ live tenants across data centers, with zero downtime

Hot-standby replication, an atomic cutover, and how "zero downtime" was proven with packet captures instead of optimism.

Read the case →
migration · failover · RCA

A one-person platform team

Running 60+ production tenants and hundreds of containers across two data centers, by architecture. Including why not Kubernetes, yet.

Read the case →
platform · GitOps · HA

12 years of homelab: my self-hosted AI stack

Local LLMs on a 24 GB GPU, hybrid RAG, a personal knowledge graph, and an autonomous agent that pentests my own SaaS.

Read the case →
LLMOps · RAG · zero-trust

A whole AI stack on one used RTX 3090

A ternary 27B model with vision, document search and image generation on one 24 GB card, a VRAM scheduler, and what broke along the way.

Read the case →
LLMOps · inference · VRAM
03

Experience

2024-presentCo-founder & CTO · MotorSyncPlatform & infrastructure, end to end, for a multi-tenant SaaS
2025-presentSenior Business Operations Manager · ContabilizeiFive operations structures, KPI-driven (+62% productivity)
2022-2025Head of Operations, Marketplace · RappiThree roles to Head. Delivery P&L, fulfillment, integrations
2021-2022IT Director · AMLRebuilt the tech org on AWS + Docker. 10+ engineers, 40+ people across four companies
2021Technical Project Manager · BairesDevFour projects for US direct-hire clients
2012-2017Project Manager · Mectron / Bradar (Embraer Defense)Radar and defense contract portfolios
2008-2011Industrial automation & data integrationOPC↔Web in Java. The original technical root
04

Stack

Platform

  • Linux · Proxmox · KVM
  • Docker · GitOps
  • Traefik · Cloudflare
  • GitHub Actions

Data & HA

  • PostgreSQL · Patroni
  • MongoDB replica sets
  • Redis Sentinel
  • PgBouncer · ZFS · NVMe

Observability

  • Prometheus · Grafana
  • VictoriaMetrics
  • Fluent Bit
  • Blackbox probing

AI infra

  • llama.cpp · Qwen
  • Embeddings · RAG
  • pgvector
  • GPU inference

Runtime I operate

  • Node.js · NestJS
  • Next.js
  • Bull queues

Identity & security

  • OAuth 2.0 / OIDC · JWT · RBAC
  • Zero-trust (Tailscale) · Cloudflare Access
  • Cloudflare Tunnel
  • Secrets mgmt · least privilege
  • Semgrep · Trivy · Grype in CI
05

Contact

Always happy to compare notes with people who run real infrastructure.