Senior AI systems software engineer focused on production ML systems, LLM adaptation and evaluation, RF/sensor intelligence, and GPU-ready inference pipelines.
I build the software layer around models: data pipelines, preprocessing, training/evaluation harnesses, inference services, telemetry, containerized deployment, and Linux systems that make ML workloads repeatable outside a notebook.
- Built a production-style AI inference service with FastAPI, ONNX-ready runtime loading, deterministic fallback inference, Prometheus metrics, Docker, benchmarks, tests, and CI.
- Built LLM adaptation and evaluation workflows covering LoRA, QLoRA, SFT, activation-level unlearning, prompt perturbation analysis, and reasoning benchmarks.
- Built RF/IQ ML pipelines with live SDR replay/receive, ONNX export, TensorRT deployment on NVIDIA Jetson, and Nsight-profiled edge inference.
- Built data and ML workflows using Kafka, S3, Airflow, TensorFlow, Docker, and AWS.
- LLM adaptation and evaluation: LoRA, QLoRA, SFT, activation-level unlearning, reasoning benchmarks
- AI inference systems: preprocessing, model execution, postprocessing, observability, reproducible deployment
- RF and sensor ML: IQ data, spectrograms, modulation recognition, SDR capture, signal-intelligence workflows
- GPU-ready experimentation: PyTorch, TensorFlow, CUDA-enabled environments, Docker, Apptainer, Slurm
- Production software: Python/C++ services, FastAPI, WebSocket streaming, CI/CD, Linux deployment
gpu-inference-pipeline
Production-style AI inference service with FastAPI, ONNX-ready runtime loading, deterministic fallback inference, Prometheus metrics, Docker deployment, benchmark tooling, tests, and CI.
masked-emotion-lora-benchmark
Reproducible LoRA adaptation benchmark for masked-emotion reasoning across encoder and generative model families. Includes continuation training, centralized metrics, and comparison tables showing RoBERTa-large improving from a paper baseline near 0.377 AccV to 0.8304 AccV.
qwen-sft-reasoning-benchmark
End-to-end supervised fine-tuning workflow for Qwen2.5-3B-Instruct using LLaMA-Factory, Hugging Face tooling, benchmark evaluation, and reasoning-task analysis.
llm-preference-unlearning
Research engineering prototype for activation-guided masked LoRA, concept suppression, prompt perturbation analysis, Fisher/saliency profiling, and before/after behavior evaluation.
rf-signal-intelligence
GPU-ready RF/IQ machine-learning workspace for modulation recognition, radar waveform analysis, signal preprocessing, recurrent/CNN hybrid models, Docker/Apptainer environments, and benchmark artifacts across RML2016, RML2018, and DeepRadar2022.
sdr-shark
Applied signal-intelligence platform with React UI, Python backend, SDR streaming, signal statistics, annotation tools, and a path toward ML-assisted RF classification.
rf-sentinel
Multi-protocol RF intelligence platform for passive discovery, normalized event processing, SDR gateway integration, Bluetooth/BLE/Zigbee/Sub-GHz plugins, dashboards, alerts, and reports.
climate-trend-predictor
Kafka/S3/Airflow/TensorFlow data pipeline for time-series ML workflow experimentation, batch and stream ingestion, visualization, and model training.
multi-agent-market-intelligence-system
Multi-agent intelligence workflow for structured research, document processing, and decision support.
Python, C++, PyTorch, TensorFlow, Hugging Face Transformers, LLaMA-Factory, ONNX, CUDA, OpenCV, FastAPI, Flask, React, Docker, Apptainer/Singularity, Slurm, Kafka, Airflow, AWS S3/EC2, Linux, SDR/SoapySDR, Jetson-class edge deployment.
LinkedIn: https://www.linkedin.com/in/rameyjm
Hugging Face: https://huggingface.co/rameyjm7

