2nd Place — Medpace Track
→ project: Pulse
Recognized by Medpace clinical-research judges for a patient-safety ecosystem combining biometric anomaly detection with conversational symptom reporting and automated MedDRA/CTCAE grading.

Harikeshav Rameshkumar
Systems & AI/LLM Engineer
I build fast, low-level systems and production AI.
>a bit more about me
I'm a computer science student at The Ohio State University who likes the hard parts of software: distributed systems, compilers, cryptography, and shipping AI that actually holds up in production.
This past year I've cut LLM extraction latency by 74% on GE Aerospace's financial pipelines, hand-written AVX2/NEON kernels for a distributed inference engine that runs Llama 70B across consumer hardware, and built a fully-homomorphic-encryption ML runtime in Rust that keeps inputs encrypted end-to-end.
When I'm not shipping production code I'm usually at a hackathon — I've placed at RevolutionUC, TartanHacks, NextHacks, and HackOHI/O with teammates I trust, building everything from clinical-trial safety platforms to gamified finance twins.
Digital Technology Intern
Shipped production LLM contract-extraction pipelines on AWS Bedrock + Claude, financial-lineage AI systems on Neptune, and adoption analytics in Polars/DuckDB across GE Aerospace's financial platform.
AI Automation Apprenticeship
Built automated AI research and competitive intelligence agents in n8n and delivered a Market Intelligence Dashboard with automated pipelines for the Rugs team.
Technical Fellow
Built end-to-end data workflows and interactive data applications in Palantir Foundry, creating ingestion pipelines, Ontology objects, and visualizations.
Software Engineering Intern
Built enterprise RAG assistants, fine-tuned BERT triage engines, and high-throughput vLLM microservices on AWS EKS.
Research Intern
Researched physics-informed machine learning for indoor wireless signal propagation, predicting indoor Wi-Fi signal strength to replace manual site surveys.
Encrypted ML inference engine (Rust + Python)
A privacy-preserving ML inference library that runs PyTorch / sklearn / XGBoost models under Fully Homomorphic Encryption — inputs and outputs never leave ciphertext. A 'narrow-waist' 3-layer architecture lowers ONNX graphs to a versioned IR executed on tfhe-rs primitives, with a PyO3 bridge and bit-for-bit exactness guarantees.
Distributed LLM inference engine in C++20
A distributed LLM inference engine that runs models like Llama 3 70B across a ring of heterogeneous consumer devices using pipeline parallelism — bypassing the VRAM wall. Hand-written AVX2/NEON SIMD kernels, a zero-copy Linux kernel-module transport, and a custom INT8/FP32 weight serializer.
Twin-fleet incident response agent with formal safety verification
An incident-response agent that rehearses candidate remediations across isolated Kubernetes twin environments receiving live-mirrored traffic before touching production. A deterministic rubric scores candidates in parallel, while a Z3-backed SMT safety kernel formally verifies system invariants to veto unsafe actions.
Intelligent LLM context-compression engine
A high-performance LLM input-compression framework. A fine-tuned BERT token classifier scores every token by semantic necessity, then a two-tier pruning engine (context + token level) cuts prompt size while a zero-hallucination constraint engine preserves code and structure.
Low-latency prediction-market logical arbitrage engine
A low-latency algorithmic trading system that detects and exploits mathematical inconsistencies across 1,000 live Polymarket order books in ~14ms. An LLM reasoning core (watsonx.ai + LangGraph) maps logical relationships across prediction markets, coupled with a Rust execution engine featuring Fill-or-Kill settlement, circuit breakers, and fixed-point math.
Local-first terminal-native AI job-application co-pilot
A local-first, terminal-native job-search platform featuring an interactive Textual TUI, Typer CLI, and background daemon over SQLite (WAL). A pluggable provider abstraction orchestrates Claude Code, Codex, and LiteLLM models with a 5-rung recovery ladder, truth-anchored resume tailoring, and a 100% line and branch test coverage gate.
→ project: Pulse
Recognized by Medpace clinical-research judges for a patient-safety ecosystem combining biometric anomaly detection with conversational symptom reporting and automated MedDRA/CTCAE grading.
→ project: Penny
Praised by Visa technical leads for uniting multimodal receipt parsing, historical spending vector search, and a real-time 'true cost' e-commerce browser extension.
→ project: Distill
Awarded for high-efficiency LLM infrastructure — a 68% token reduction and 37% latency drop while maintaining 99% accuracy across massive context windows.
→ project: LeadForge
Honored at OSU's flagship competition for agentic AI architecture — multi-agent lead scoring, autonomous React prototype generation, and real-time bi-directional AI voice calling.
→ project: VR Market Simulator
Top game-dev honors from Meta and OSU Linguistics for immersive language learning — Meta Quest hand-tracking, dynamic VR budgeting, and NavMesh crowd AI.
Repeated recognition for rapid full-stack execution and cryptographic security implementation under tight 24–48 hour competitive sprints.
Awarded Dean's List and University Honors every semester while maintaining a 3.9 / 4.0 cumulative GPA in Computer Science.
>got a role, a project, or just want to say hi? my inbox is open.
$whoami
Harikeshav Rameshkumar