Hi, I'm
Kazusa Tsubota
Senior AI Platform Engineer
Cloud infrastructure, AI/ML production, and backend engineering — designing and operating vLLM-based inference on Kubernetes/EKS, multi-model serving, AI Gateway routing, and production LLM services from RAG to tool calling.
- AWS
- EKS
- vLLM
- Kubernetes
- Python
About
Senior AI Platform Engineer with 7+ years of software engineering experience specializing in production AI infrastructure, LLM inference, GPU-based model serving, and Kubernetes/EKS.
Hands-on experience designing and operating vLLM-based inference infrastructure, multi-model serving and routing, AI Gateway architecture, autoscaling, Terraform, Docker, and CI/CD — with a strong focus on inference performance, reliability, observability, security, and cloud cost optimization.
Experienced in connecting AI infrastructure with production application architecture, including LLM-powered services, embeddings, semantic search, pgvector, RAG, and LLM tool calling.
Experience
-
ScalyX.ai
AWS AI Infrastructure / AI Production Engineer
- Architected and migrated a multi-tenant AI-powered retail platform from per-tenant EC2 to a shared Kubernetes platform on AWS EKS, consolidating workloads across 30+ tenants and enabling demand-based autoscaling while reducing average monthly infrastructure costs by approximately 17%.
- Designed and optimized vLLM-based GPU inference infrastructure on Kubernetes/EKS, evaluating model characteristics, GPU requirements, workload patterns, and serving configurations to balance inference latency, throughput, availability, and cost.
- Implemented an AI Gateway and multi-model routing layer across LLM, speech, and vision workloads, routing requests based on workload and resource requirements and reducing GPU-related infrastructure costs by 15% while maintaining target service performance.
- Built production AI services and workflows using Python/FastAPI, PostgreSQL, Docker, embeddings, semantic search, pgvector, RAG, and LLM tool calling, connecting retrieval, generation, and controlled backend actions.
- Designed scalable Kubernetes workloads using ALB, health checks, HPA, cluster/node autoscaling, and Helm, while standardizing infrastructure deployment with Terraform, Docker, ECR, and CI/CD.
- Built end-to-end observability with Prometheus, Grafana, OpenTelemetry, Jaeger, and CloudWatch, reducing average incident investigation time from 50 minutes to 30 minutes (40% reduction).
- Applied AWS/Kubernetes security and tenant-isolation controls using IAM, VPC/private networking, Security Groups, Kubernetes RBAC, Secrets Manager, and tenant-level access controls; performed synthetic load testing that validated aggregate request throughput of up to 10K requests/sec under controlled load while evaluating latency, concurrency, and GPU utilization.
-
Madoromi, Inc.
Infrastructure / Backend Engineer
- Consolidated six backend services from six dedicated EC2 instances into a shared t3.medium environment using NestJS monorepo and Turborepo, preserving logical service boundaries while reducing the production EC2 footprint by 83% (6→1) and monthly EC2 cost from ~$45 to ~$30 (33%) for Alpha/MVP traffic.
- Designed and operated AWS production infrastructure across EC2 and Aurora RDS, right-sizing compute and database capacity against observed workload patterns rather than maintaining independent per-service capacity.
- Established a modular TypeScript/NestJS backend architecture with PostgreSQL and shared domain/API boundaries, collaborating with frontend engineers on feature specifications, API contracts, and implementation requirements while structuring services for future containerization and Kubernetes/EKS adoption.
- Automated non-production infrastructure scheduling with AWS Lambda and EventBridge, reducing development-environment runtime by approximately 70% (~50 vs. ~168 hours/week per instance) while preserving availability during engineering hours.
-
StoreHub
Full-Stack Engineer
- Built commerce and POS capabilities using Next.js, TypeScript, React, Node.js, REST APIs, and PostgreSQL, supporting transaction, catalog, inventory, customer, and multi-location merchant workflows across a cloud SaaS platform.
- Designed reliable distributed and asynchronous systems using idempotent APIs, Redis Streams consumer groups, background processing, and business-rule validation for e-commerce, POS, inventory synchronization, order processing, and notifications.
- Optimized high-traffic product and inventory services with tenant-scoped Redis caching and PostgreSQL query/index/schema optimization, achieving a 90%+ cache hit rate and processing 1,000+ notifications/sec with <150ms p95 delivery latency.
Skills
AI Infrastructure & Platform
LLM inference & GPU infrastructure · vLLM · multi-model serving & routing · AI Gateway · inference optimization · Kubernetes · EKS · HPA · Helm · Cluster & node autoscaling
AI Engineering
RAG · embeddings · semantic search · PostgreSQL / pgvector · LLM tool calling · AI agent architecture
Cloud & Infrastructure
AWS · EC2 · EKS · ALB · VPC · RDS / Aurora · S3 · ECR · CloudWatch · IAM · cost optimization · right-sizing
IaC, DevOps & Observability
Terraform · Docker · CI/CD · Prometheus · Grafana · OpenTelemetry · Jaeger · Metrics · logs · distributed tracing
Backend & Security
Python · FastAPI · TypeScript · NestJS · PostgreSQL · Redis Streams · REST APIs · Kubernetes RBAC · tenant isolation · secrets management
Education
Kobe University
Bachelor of Engineering in Computer and Intelligent Systems Engineering
2018
Contact
Open to senior AI platform, ML infrastructure, and production backend engineering opportunities.
Japan · Remote
- Email kazusa931209@gmail.com
- LinkedIn linkedin.com/in/kazusa-tsubota
- GitHub github.com/tsubokazu