Professional headshot of Kazusa Tsubota in a dark suit and white shirt

Hi, I'm

Kazusa Tsubota

Senior AI Platform Engineer

Cloud infrastructure, AI/ML production, and backend engineering — designing and operating vLLM-based inference on Kubernetes/EKS, multi-model serving, AI Gateway routing, and production LLM services from RAG to tool calling.

  • AWS
  • EKS
  • vLLM
  • Kubernetes
  • Python

About

Senior AI Platform Engineer with 7+ years of software engineering experience specializing in production AI infrastructure, LLM inference, GPU-based model serving, and Kubernetes/EKS.

Hands-on experience designing and operating vLLM-based inference infrastructure, multi-model serving and routing, AI Gateway architecture, autoscaling, Terraform, Docker, and CI/CD — with a strong focus on inference performance, reliability, observability, security, and cloud cost optimization.

Experienced in connecting AI infrastructure with production application architecture, including LLM-powered services, embeddings, semantic search, pgvector, RAG, and LLM tool calling.

Experience

  1. ScalyX.ai

    AWS AI Infrastructure / AI Production Engineer

    Jan 2025 – Present

    USA · Remote

    • Architected and migrated a multi-tenant AI-powered retail platform from per-tenant EC2 to a shared Kubernetes platform on AWS EKS, consolidating workloads across 30+ tenants and enabling demand-based autoscaling while reducing average monthly infrastructure costs by approximately 17%.
    • Designed and optimized vLLM-based GPU inference infrastructure on Kubernetes/EKS, evaluating model characteristics, GPU requirements, workload patterns, and serving configurations to balance inference latency, throughput, availability, and cost.
    • Implemented an AI Gateway and multi-model routing layer across LLM, speech, and vision workloads, routing requests based on workload and resource requirements and reducing GPU-related infrastructure costs by 15% while maintaining target service performance.
    • Built production AI services and workflows using Python/FastAPI, PostgreSQL, Docker, embeddings, semantic search, pgvector, RAG, and LLM tool calling, connecting retrieval, generation, and controlled backend actions.
    • Designed scalable Kubernetes workloads using ALB, health checks, HPA, cluster/node autoscaling, and Helm, while standardizing infrastructure deployment with Terraform, Docker, ECR, and CI/CD.
    • Built end-to-end observability with Prometheus, Grafana, OpenTelemetry, Jaeger, and CloudWatch, reducing average incident investigation time from 50 minutes to 30 minutes (40% reduction).
    • Applied AWS/Kubernetes security and tenant-isolation controls using IAM, VPC/private networking, Security Groups, Kubernetes RBAC, Secrets Manager, and tenant-level access controls; performed synthetic load testing that validated aggregate request throughput of up to 10K requests/sec under controlled load while evaluating latency, concurrency, and GPU utilization.
  2. Madoromi, Inc.

    Infrastructure / Backend Engineer

    Jan 2022 – Dec 2024

    Tokyo, Japan · Hybrid

    • Consolidated six backend services from six dedicated EC2 instances into a shared t3.medium environment using NestJS monorepo and Turborepo, preserving logical service boundaries while reducing the production EC2 footprint by 83% (6→1) and monthly EC2 cost from ~$45 to ~$30 (33%) for Alpha/MVP traffic.
    • Designed and operated AWS production infrastructure across EC2 and Aurora RDS, right-sizing compute and database capacity against observed workload patterns rather than maintaining independent per-service capacity.
    • Established a modular TypeScript/NestJS backend architecture with PostgreSQL and shared domain/API boundaries, collaborating with frontend engineers on feature specifications, API contracts, and implementation requirements while structuring services for future containerization and Kubernetes/EKS adoption.
    • Automated non-production infrastructure scheduling with AWS Lambda and EventBridge, reducing development-environment runtime by approximately 70% (~50 vs. ~168 hours/week per instance) while preserving availability during engineering hours.
  3. StoreHub

    Full-Stack Engineer

    Dec 2018 – Dec 2021

    Malaysia · Remote

    • Built commerce and POS capabilities using Next.js, TypeScript, React, Node.js, REST APIs, and PostgreSQL, supporting transaction, catalog, inventory, customer, and multi-location merchant workflows across a cloud SaaS platform.
    • Designed reliable distributed and asynchronous systems using idempotent APIs, Redis Streams consumer groups, background processing, and business-rule validation for e-commerce, POS, inventory synchronization, order processing, and notifications.
    • Optimized high-traffic product and inventory services with tenant-scoped Redis caching and PostgreSQL query/index/schema optimization, achieving a 90%+ cache hit rate and processing 1,000+ notifications/sec with <150ms p95 delivery latency.

Skills

AI Infrastructure & Platform

LLM inference & GPU infrastructure · vLLM · multi-model serving & routing · AI Gateway · inference optimization · Kubernetes · EKS · HPA · Helm · Cluster & node autoscaling

AI Engineering

RAG · embeddings · semantic search · PostgreSQL / pgvector · LLM tool calling · AI agent architecture

Cloud & Infrastructure

AWS · EC2 · EKS · ALB · VPC · RDS / Aurora · S3 · ECR · CloudWatch · IAM · cost optimization · right-sizing

IaC, DevOps & Observability

Terraform · Docker · CI/CD · Prometheus · Grafana · OpenTelemetry · Jaeger · Metrics · logs · distributed tracing

Backend & Security

Python · FastAPI · TypeScript · NestJS · PostgreSQL · Redis Streams · REST APIs · Kubernetes RBAC · tenant isolation · secrets management

Education

Kobe University

Bachelor of Engineering in Computer and Intelligent Systems Engineering

2018

Contact

Open to senior AI platform, ML infrastructure, and production backend engineering opportunities.

Japan · Remote