a
amitkumar_7083

Amitkumar J

@amitkumar_7083

Cloud DevOps Engineer

India
Inglese
Alcune informazioni sono riportate in lingua inglese.
Chi sono
DevOps Expert, Excellent at GCP cloud IAC terraform, Ansible. Hands-on Experience of Patching and Security of Infra used in application for complied. Strong Linux Admin skills. Deployed and Maintained Grafana Monitoring Stack using Helm on Kubernetes, GKE/OpenShifts. Azure DevOps for deployement of Dev, Ppd and Prod Environment on K8S. used Istio Virtual Service and Gateway for routing. Strong Knowledge of ELK stack Including Kafka used on large Scale Microservice.... Continua a leggere

Competenze

a
amitkumar_7083
Amitkumar J
offline • 
Tempo di risposta medio: 1 ora

Consulta i miei servizi

Containerizzazione DevOps
I will devops and cloud solutions with grafana monitoring

Portfolio

Esperienza lavorativa

IBM

Cloud DevOps Engineer

IBM • Full time

Oct 2024 - Present1 yr 10 mos

• Cloud Storage Cluster Rack/ Stack Configurations in platform-inventory files using automations • Setup Storage cluster of Ceph (SDS) using ansible-Jenkins, by continuous deployment. Providing bugs/troubleshooting fix to the automations team. Bringup include Firmware, Encryption Keys for drives, OS installation and Ceph platform configurations. • Implemented continuous release upgrades on storage cluster using Jenkins CD Jobs. Include Firmware , OS Upgrade and Ceph Upgrades. • Executed load and performance benchmarking for Ceph Storage Block Volumes and CephFS File Shares to validate IOPS, throughput, latency, and storage reliability before production rollout. Added Validation methods. • Conducted NVMe drive soak testing and performance validation on newly installed storage servers to identify hardware defects prior to Ceph cluster deployment. • Detected and resolved Ceph MON quorum loss by introducing a temporary MON on a healthy OSD node, restoring quorum before replacing faulty MONs; validated across cluster bootstrap scenarios. • Tested maintenance and recovery operations for Ceph OSD, MON and NFS nodes, including graceful node removal, hardware replacement, cluster reintegration ensuring cluster stability. • Hardware remediation on cluster nodes, including component replacement (CPU, RAM, NICs, Drives, Motherboards, PSUs), OS redeployment using CD, and seamless addition of nodes into the cluster. Documented the Steps as SOP for operations work. • Provided production support & on-call operations for IaaS Storage, troubleshooting Ceph clusters and infrastructure security components. Diagnosed network and switch-level issues in collaboration with networking teams to restore service availability. • Built CD jobs for Sanity of client I/O workload Test during Storage Cluster Upgrade in dev, stage and prod environments. • Optimized CD jobs to automatically unlock storage node drives after reboot. • Saved 3 to 4 engineering hours daily by automating alert snoozing dur

Reliance_Jio Infocomm

DevOps Engineer

Reliance Jio Infocomm • Full time

Sep 2022 - Aug 20241 yr 11 mos

• Design and Implemented ScyllaDB -> Kafka -> Elasticsearch, data pipeline with kafka source and sink connectors and benchmark the Traffic. • Designed/Built Stable Data Pipeline from MySql CDC -> Kafka -> ELK & ScyllaDB CDC -> Kafka -> ELK, both Data merge same index in Elasticsearch Optimized Queries performance by 98% render data on UI. • Designed scaling automation on connect server to increase the filtering process using golang code instances, To handle traffic upto 4-5 Cr data published into kafka. Automated ScyllaDB CDC source connector for daily table creation using python scripts removed Manual Efforts of one person. • Designed Load test methodology to setup benchmark for Data-pipeline also Load/Stress test on Scylla DB using Cassandra Stress and Golang script using goroutine concurrency. • Designed, deployed, and managed production Kubernetes clusters (GKE, Self-Managed with 100+ microservices) Cilium, HAProxy, and Istio, Implementing HPA autoscaling, kubeadm certificate renewal, and end-to-end workload management. • Automated infrastructure provisioning using Terraform and Ansible for 400+ cloud resources, improving consistency and reducing provisioning time by 60%. • Built CICD Jobs in azureDevOps for deployment strategies including rolling upgrades, rollback mechanisms for Grafana Stack using helm charts. • Integrated security and code quality checks (Fortify, SonarQube, Black Duck) into pipelines, reducing vulnerabilities by 90 and ensuring compliance. • Enabled secure and resilient systems by enforcing RBAC, security policies, and container/OS hardening, reducing vulnerabilities by 65%. • Setup and Managed ScyllaDb, Kafka and ElasticSearch for handling the 4-5Cr of transactions. Elasticsearch node removals and additions. • Deployed and tested HA Postgresql Master failover using repmgr tool on kubernetes using Helm charts. Backup and restore of Postgresql for Grafana UI as recovery in case of Repmgr failover failures. • IAM Access management, ServiceAc