H: 100VH
W: 100%
System Architecture
VISHAL GUNJAL

Architecting
Resilient Systems .

Orchestrating Cloud-Native & Agentic AI Infrastructure.

I don't just configure tools; I engineer systems. Bridging the gap between application logic and cloud-native infrastructure through first-principles Systems Thinking. Currently pioneering Agentic Workflows on Amazon Bedrock.

Core Stack: Terraform Kubernetes AWS Jenkins Go
14+
Production Repos
25+
DevOps/AI Tools
5+
AWS Services
7K+
Network
Vishal Gunjal - SRE Professional Headshot
BeSA - Agentic AI · 2026
DevOps / SRE / Cloud Engineer
Grafana GitOps DevSecOps Go - Python Jenkins SonarQube ELK Stack MCP Vercel Edge Network Kubernetes Terraform Argo CD Amazon Bedrock Agentic AI Prometheus Grafana GitOps DevSecOps Go - Python Jenkins SonarQube ELK Stack MCP Vercel Edge Network Kubernetes Terraform Argo CD Amazon Bedrock Agentic AI Prometheus
H: AUTO
SEC-01
VISHALOS / ENGINEERING / CORE / ABOUT
# ABOUT —

Building the Future of Cloud & AI
Platforms

COM_Rpx_GRID

"If a system can't tell you what it's doing, it's already broken—you just haven't been paged yet."

I'm Vishal, an SRE-focused engineer working on the infrastructure layer that sits between cloud platforms and the teams who build on top of them. I focus on DORA metrics and building systems that improve reliability and developer experience.

More recently, I’ve been exploring the intersection of Infrastructure and AI, experimenting with agentic workflows on Amazon Bedrock using the Model Context Protocol (MCP), and building assistive systems where LLM-based tools can help analyze failures and surface actionable insights.

MISSION_STATEMENT

Champion of reliability.
Builder of resilient systems.

I strongly believe that reliability is a feature. My focus is on building platforms that are easy to operate, observable, and practical to scale - while improving lead time for changes and time to detect issues through continuous iteration and real-world testing.

14+
DevOps Projects
3+
Cloud Platforms
20+
DevOps Tools
AI+Cloud
Engineering Focus

Cloud Platform Engineering

Architecting multi-AZ EKS clusters with modular Terraform, declarative ArgoCD GitOps, and Istio service mesh. Infrastructure that provisions itself, heals itself, and documents its own state.

SRE & Reliability

Treating uptime as a first-class engineering requirement. Full observability stacks, error budget tracking, runbook-driven incident response, and DevSecOps pipelines with zero-CVE deployment gates.

Agentic AI Infrastructure

Building the next layer of ops: LLM-powered agents on Amazon Bedrock that analyze build failures, correlate distributed traces, and orchestrate infrastructure changes via the Model Context Protocol.

Location & Languages

Based in Pune, India. English · Hindi · Marathi. Open to remote collaboration on cloud, DevOps, and AI infrastructure projects globally.

AGENT_TRACE: BEDROCK_MCP_INVOCATION
> Initializing Amazon Bedrock Agent (Claude-3-Opus)...
> Context: Incoming PagerDuty alert - EKS Cluster AP-South-1
> Invoking Model Context Protocol (MCP) tool: analyze_traces
[WARN] High latency detected in 'payment-service' container
> Correlating traces with CloudWatch metrics...
[INFO] Root cause identified: RDS connection pool exhaustion.
> Formulating remediation strategy -> Update Terraform connection_limit parameter.
> AI Agent: "Drafting PR for terraform-aws-rds module..."
VishalOS / Toolbox / SRE Stack / Capabilities
Technical Expertise

Skills & Technologies

Deep expertise across the full cloud-native stack - from bare-metal infrastructure to AI-driven agentic systems.

CORE Engineering Foundations

Systems-first engineering grounded in OS internals, DBMS architecture, concurrency models, and protocol-level networking.

Linux Kernel / OS DBMS (Postgres / Redis) TCP/IP Networking Java / Go / Python LLD & Concurrency Distributed Systems Bash / Shell Scripting System Design
SYSTEM-LEVEL ANALYSIS
DETERMINISTIC THINKING

SRE & Reliability Engineering

Engineering resilient distributed systems through observability, failure analysis, tracing, and strict SLO governance.

Failure Analysis Distributed Tracing Prometheus / Grafana ELK Observability SLIs / SLOs MTTR Reduction Incident Response High Availability
99.999% AVAILABILITY GOAL

Cloud-Native Platform Engineering

Architecting scalable, self-healing infrastructure platforms using Kubernetes, GitOps, Terraform, and distributed cloud-native patterns.

Kubernetes (EKS) Terraform (IaC) ArgoCD (GitOps) Docker / Helm AWS Ecosystem Ansible RabbitMQ Multi-Region Arch
MULTI-REGION PLATFORM SCALE

DevSecOps & Governance

Embedding security, policy enforcement, and compliance automation directly into cloud-native delivery pipelines.

Jenkins / GH Actions SonarQube (SAST) Trivy Security OPA Gatekeeper IAM / IRSA Blue-Green / Canary Policy Enforcement
SHIFT-LEFT SECURITY

AI & Agentic Infrastructure

Designing GenAI infrastructure and intelligent workflow orchestration through MCP-driven agentic systems.

Amazon Bedrock Agentic Workflows MCP Architecture AI-Ops Integration LLM APIs Context Engineering Toil Automation
AGENT-DRIVEN AUTOMATION

Systems Thinking Approach (STA)

Applying first-principles engineering to reliability, scalability, distributed systems behavior, and infrastructure decision-making.

First-Principles Systems Analysis Performance Modeling Scalability Engineering FinOps / Cost Optimization Infrastructure Tradeoffs
TOOLS CHANGE • FUNDAMENTALS REMAIN
VISHALOS / ARCHITECTURE / TELEMETRY
# STACK ARCHITECTURE —

Infrastructure & Observability

COM_Rpx_GRID

A live view into the stack powering this platform. Built on cloud-native principles, automated by code, and monitored in real-time.

AWS EKS Cluster
HEALTHY
Active Nodes 12
Running Pods 148
Control Plane v1.28
Terraform State
IN_SYNC
$ terraform plan
Acquiring state lock...
Refreshing state...
No changes. Your infrastructure matches the configuration.
Plan: 0 to add, 0 to change, 0 to destroy.
Observability
LIVE
p99 Latency 42ms
CPU Load 45%
/ CI/CD Pipeline
SYNCED
Build
Test
Image
Argo Sync
#8f2d91a — feat: agentic control plane
VishalOS / SRE AWS History / Systems Lifecycle
Systems Engineering Timeline

Systems Engineering & Operational Lifecycle

Architecting resilience across distributed systems. From scaling Kubernetes to deploying Agentic AI, every role is a chapter in building the 'Invisible Infrastructure'.

# CREDENTIALS —

Verified Credentials

COM_Rpx_GRID

Industry-recognized certifications and ongoing learning paths in cloud-native architecture.

05
VISHALOS / ACADEMIC / BACKGROUND

Education

Academic foundations that shaped my engineering mindset and technical depth.

Community

Community Roles

Contributing to cloud-native ecosystems while architecting the next generation of reliable cloud and AI infrastructure.

Google Developer Group

Actively involved in Pune’s developer ecosystem through GDG, engaging in deep technical discussions around cloud-native systems, distributed architectures, and modern backend patterns.

AWS User Group Pune

An active contributor to the AWS community, engaging in real-world discussions around scalable cloud architecture, DevOps, and production-grade systems. Exploring Agentic AI workflows with Amazon Bedrock.

CNCF Community

Deeply aligned with the cloud-native ecosystem, working with Kubernetes, observability, and platform engineering. Implementing CNCF tools in real-world projects and GitOps workflows.

FinOps Community

Engaged in FinOps practices to bridge engineering with cost efficiency. Optimizing cloud spend across AWS and Kubernetes workloads for operational clarity.

Atlassian Community

Refining workflow automation and team collaboration. Leveraging Atlassian Rovo and AI to enhance knowledge discovery and DevOps lifecycle management.

Snowflake Community

Exploring modern data platforms and cloud data warehousing. Connecting backend systems and DevOps practices with evolving data engineering workflows.

“Learning, building, and growing through active community engagement in modern engineering.”
VISHALOS NETWORK ECOSYSTEM ENGAGEMENT

Community Engagement & Events

Participating in cloud-native meetups, workshops, and developer events while continuously learning in the DevOps, Cloud, and AI ecosystem.

TECH MEETUPS ATTENDED

60+

GLOBAL/APAC CONFERENCES

25+

CLOUD COMMUNITIES

3+

ACTIVE SINCE

2022
SYSTEM ONLINE
NODE // SRE-PROOF-01
AWS / Cloud
AWS Summit Mumbai, India
NODE // SRE-PROOF-02
AWS / Cloud
Kubernetes Day Badge
NODE // SRE-PROOF-03
AWS / PLATFORM
AWS Summit Badge
NODE // SRE-PROOF-07
CNCF / COMMUNITY
Kubernetes Meetup Panel
NODE // SRE-PROOF-08
SRE / PANELIST
SRE Speaker Panel
NODE // SRE-PROOF-09
GDG / DEVFEST
GDG DevFest Pune 2025
NODE // SRE-PROOF-06
DEVOPS / MIXER
DevOps Ecosystem Mixer
NODE // SRE-PROOF-05
GOOGLE / GUEST
Google Guest Badge
SRE in Practice

Live Mission Control Dashboard

A high-fidelity simulation of a production cloud environment. Track Golden Signals, monitor GitOps deployment streams, and manage error budgets in real-time. Syncs directly with the SRE Terminal.

EKS Status: NOMINAL
142ms
Latency (p99)
Within SLO
0.08%
Error Rate
Healthy
1847rps
Traffic
Total throughput
64%
Saturation (EKS)
Healthy
94%
Cloud Efficiency
Optimal
$0.004/1k req
Cloud ROI
82% Spot Util

IaC Blueprint

VIEW_ARCHITECTURE
Budget Remaining 99.99%
Latency Trends
Traffic Load

Operational Log

LIVE_FEED
Cluster: cluster-prod-01
GitOps: ArgoCD (Active)
Engineering Principles

System Design Philosophy

How I approach building systems that work reliably in the real world - shaped by hands-on experience across projects and production environments.

Principle 01
Observability First, Always

I don’t treat observability as an add-on. Before shipping anything, I make sure the system can explain what it’s doing. If it can’t, debugging becomes guesswork. Metrics, logs, and traces help me understand system behavior early - so when something breaks (and it will), I’m not starting from zero.

Principle 02
Design for Failure, Not Just Success

In distributed systems, failure is normal - not an edge case. I design systems with that in mind from the beginning. Retries, circuit breakers, and graceful degradation are part of the design - ensuring the system behaves predictably even under stress.

Principle 03
Document Decisions, Not Just Code

Code shows what was built, but not why. I focus on documenting the reasoning behind decisions - trade-offs, constraints, and alternatives. This helps both teammates and my future self understand systems without rethinking everything from scratch.

Principle 04
Automation Over Repetition

If I find myself doing something more than once, I automate it. Repetition leads to inconsistency and human error. From CI/CD pipelines to infrastructure provisioning, I focus on building systems that are repeatable, predictable, and require minimal manual intervention.

Principle 05
AI-Augmented Reasoning

Leveraging Large Language Models and Agentic workflows not as a replacement, but as a force-multiplier for root cause analysis and complex system synthesis.

Principle 06
Reliability as a Habit

Reliability isn't a one-time setup; it's a daily practice. From small commits to large architectural shifts, every decision is measured against its impact on the 99.99% availability goal.

Principle 07
Fundamentals Over Frameworks (STA)

Before abstracting complexity to Kubernetes or AWS, I analyze the underlying compute, OS thread management, and raw network topology. Tools change yearly; TCP/IP, DBMS locks, and Linux cgroups do not. I apply a Systems Thinking Approach to every architectural decision.

“Applied thinking from real systems, not just theory.”
VishalOS / Engineering / Philosophy / Mindset
How I Think

How I Debug Production Systems

A structured, hypothesis-driven approach to debugging distributed systems - focusing on signals, narrowing scope, and reaching root cause without guesswork.

Step 01

Identify the Symptom

I start from the edge - typically ingress or API gateway logs (Nginx/Envoy) - to understand what users are actually experiencing. I look for clear signals like HTTP 5xx errors, latency spikes, or dropped connections, and correlate timestamps to establish when the issue started.

Step 02

Scope the Blast Radius

Before diving deeper, I determine whether the issue is isolated or systemic. Using metrics from Prometheus/Grafana, I check patterns - CPU, memory, request rates, DB connections - to see if multiple services are affected or just a single pod/node.

Step 03

Pinpoint the Bottleneck

Once scoped, I follow the request path using distributed tracing (Jaeger/Datadog). This helps identify the exact service or span causing the issue - whether it’s a slow downstream dependency, a failing database query, or unexpected latency between services.

“Debugging with signals, not assumptions.”
Case Study

DevSecOps Enterprise Pipeline

An end-to-end cloud-native platform for a 3-tier MERN application on Vercel - code commit to production with security gates at every stage. Click any stage to inspect it.

DevOps-Wanderlust-Enterprise-Pipeline
Production-grade MERN application deployed on Vercel with GitOps, full DevSecOps quality gates, and a complete observability stack. Demonstrates the full software delivery lifecycle from commit to monitored production.
EKS · Multi-AZ
DevSecOps
GitOps · ArgoCD
Cloud infrastructure architecture - multi-AZ, GitOps-driven
AWS VPC - us-east-1 Internet User traffic Route 53 DNS resolution ALB Load balancer us-east-1a us-east-1b us-east-1c Worker Node t3.medium Pod Pod Pod IRSA · NetworkPolicy Worker Node t3.medium Pod Pod Pod IRSA · NetworkPolicy Worker Node t3.medium Pod Pod Pod IRSA · NetworkPolicy EKS Control Plane - Managed by AWS (HA across 3 AZs) kube-apiserver · etcd · kube-scheduler · controller-manager ArgoCD GitOps sync API service pod Worker service pod Frontend pod
8+
Microservices
3
Security Gates
Multi-AZ
EKS Cluster
99.9%
Uptime Target
CI/CD · DevSecOps · GitOps Pipeline
All systems passing Click a stage to inspect
Git Commit
Code push to main branch triggers pipeline
Jenkins Build
Groovy pipeline · parallel stages
SonarQube
SAST · code smells · coverage gate
OWASP Check
Dependency CVE scan · NVD feed
Trivy Scan
Container image CVE · OS packages
Docker Push
ECR registry · image tag + digest
ArgoCD Sync
GitOps reconciliation · self-heal
EKS Deploy
Multi-AZ · RBAC · IRSA · Network Policy
Prometheus
RED metrics · alerting rules
SEC-03: DELIVERABLES
VishalOS / Lab / Deployments / Projects
Platform Engineering

Production Architectures & Lab Deployments

Real-world DevOps, cloud infrastructure, and AI platform projects - automation, scalable architecture, and modern cloud-native engineering.

Platform Overview
Explore All Projects
Writing

Articles & Engineering Insights

Sharing learnings on DevOps, cloud-native platforms, Kubernetes, AI infrastructure, and modern engineering practices.

Community Representation
Testimonials

What People Say

Feedback from peers, mentors, and community collaborators.

Collaboration

Working With Me

What you can expect when building critical infrastructure together - principles I ship with.

PRINCIPLE // #1

Reliability is a Feature

I don't just build pipelines - I build observability, self-healing systems, and runbooks from day one. 99.99% uptime isn't a goal, it's the baseline architecture decision.

PRINCIPLE // #2

Developer Velocity First

Infrastructure should be invisible and empowering, never a bottleneck. I measure success by how fast engineers can ship without worrying about the platform.

PRINCIPLE // #3

Security by Design

Zero-trust and compliance are integrated into the CLI and CI/CD pipeline - not bolted on after the fact. Shift-left security is standard operating procedure.

RESPONSE SLA
< 24h
Global Async Collaboration
OPTIMIZATION
5m <1m
CI/CD Pipeline Throughput
CLOUD EFFICIENCY
-40%
Avg. Infrastructure OpEx Reduction
Collaboration
Verified Partnership

Engineering Reliable Systems For A Chaotic Internet.

Open to DevOps, cloud infrastructure, platform engineering, and collaboration opportunities.

"Calm infrastructure creates fast teams."

14 + Repos 800 + Commits K8s / AWS / TF
Lighthouse: 99 TTFB: <100ms Bundle: <150KB
⚠ GRAVITY OFFLINE - type kubectl scale deployment earth-gravity --replicas=1 to restore
VishalOS / Network / Ecosystem / Community

Building Bridges in Pune's Tech Ecosystem

I believe the best infrastructure is invisible—the silent force that empowers developers and ensures a seamless experience. As an active participant in the Atlassian Community, I’m dedicated to sharing knowledge and refining the workflows that drive our industry forward.