Skip to content
View SeMe4K-0's full-sized avatar
🏠
Working from home
🏠
Working from home

Block or report SeMe4K-0

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SeMe4K-0/README.md

English Русский

Semyon Davshits - Data Scientist / ML Engineer

Resume Telegram Email

About

Fourth-year student at Bauman Moscow State Technical University (Dept. IU5), graduating in 2027. I build models on tabular data and time series, work on uplift modelling and experiment statistics, and also do computer vision and audio. I take models through to a service with an API, tests and Docker. Open to Data Science and ML internships.

Skills

Area Stack
ML · Tabular Python pandas polars scikit-learn CatBoost LightGBM NumPy SciPy statsmodels · bootstrap uplift · causal inference Optuna Jupyter Matplotlib
DL · CV · Audio PyTorch torchvision segmentation-models-pytorch albumentations Ultralytics YOLO OpenCV CoreML torchaudio · librosa Demucs · museval
LLM · NLP Transformers Gemini API CLAP · MERT · EnCodec TF-IDF · embeddings prompt engineering · evals guardrails · red team
Data · Infra SQL · PostgreSQL dbt FastAPI Docker · Compose pytest Git · Git LFS Linux Go gRPC Redis

Projects

X5 RetailHero Uplift · who actually changes behaviour when contacted  Δ uplift@10% +9.53 pp

Uplift modelling on 45.8M receipt lines in PostgreSQL (dbt, 42 tests). T-, S- and X-learners on CatBoost and the causal metrics (Qini, AUUC, uplift@k) written from scratch. The top 10% by response and the top 10% by uplift overlap by only 0.44% against 10% by chance; Δ uplift@10% = +9.53 pp [+6.63; +12.54], 2,000-replicate bootstrap. Response-based targeting loses money (0.53 pp against a 4.67 pp break-even), while uplift targeting is worth ₽612k per campaign. The analysis plan was pre-registered before the first model, and a mutation test proves the leakage auditor works.

E-CUP 2026 (Ozon) · 30-day GMV forecast for 250,000 users · team of 2  rank 14 of 316

14th of 316, private RMSLE 1.6631. Every training row is a user × date anchor with features strictly before that date, the model is trained directly in the metric's scale (log1p), and the final model is a GRU + gradient boosting ensemble. Adversarial validation exposed an anchor-date leak (AUC 1.0); once it was removed, 93% of the local improvement turned out to be illusory.

OreScope · ore grade from thin-section microscopy · Nornickel AI Science Hack finalist  macro-F1 0.91

U-Net phase segmentation on panoramas of up to 300 MP; the grade follows from an explicit expert rule over the mask. macro-F1 0.91 on intergrowth type, 0.77 end-to-end. Colour normalisation lifted the unseen camera domain from 5 to 14 samples out of 20 without new labels. A panorama takes 124 s against a 5-minute requirement; FastAPI, Docker, built in 48 hours.

Stem Separator · splitting a track into four stems  SDR vocals 9.34 dB

Demucs (htdemucs) behind a REST API, a web UI and a CLI. SDR 9.34 dB on vocals and 9.90 dB on bass (museval, MUSDB18-sample, 5 tracks). A 3:17 track is separated in 14 s on an M4 Pro (×14 real time). MPS/CUDA/CPU autodetect, chunked inference, 19 tests, Docker Compose.

FAD Benchmark · evaluating generative music models · research at Bauman MSTU  benchmark 5 models · 3 embedders

Five open-source text-to-music models (MusicGen, AudioLDM-M/L, MusicLDM, Riffusion) compared with FAD, FAD-inf, per-song FAD and CLAP Score, using three embedders (CLAP-LAION-Music, MERT-v1-95M, EnCodec) and two reference sets. The leader depends on the reference: AudioLDM-M on FMA-Pop (CLAP-FAD 0.039), MusicLDM on MTG-Jamendo (0.0044).

AI Team Assistant · LLM assistant with a strict answer format  demo live

Google Gemini behind a server-side route with structured output through responseSchema, so the JSON is valid by construction. Prompt-evals run 4 synthetic requests, including a prompt injection, and check the structure and content invariants.

Smoking Detection · detecting smoking in a video stream, on-device  mAP50 0.713

YOLO26n on person / smoke classes: mAP50 0.713, mAP50-95 0.317, precision 0.744, recall 0.680. Smoking is flagged by box overlap, and the model is exported to CoreML (.mlpackage) for inference on iOS.

View all projects

Competitions

Year Competition Result Link
2026 E-CUP 2026 (Ozon): 30-day user GMV forecast · team O3 14th of 316 RMSLE 1.6631 solution
2026 Yandex ML Challenge (Young & Yandex), final 42nd of 100 finalists certificate
2026 Yandex School of Data Analysis: AI Agents Security Week 99.81 / 100 guardrail, from 61.54 · red team 71.43 / 100 certificate
2026 DatsSol hackathon (DatsTeam): colony-control bot 16th of 166 solution
2026 Nornickel AI Science Hack: ore classification from microscopy finalist macro-F1 0.91 solution

Education

Institution Programme Year
Bauman Moscow State Technical University BSc, Informatics and Computer Engineering (09.03.01), Dept. IU5 4th year, graduating 2027
Training log: validation loss starts rising, early stopping

Pinned Loading

  1. Ozon-ecup-2026-user-ltv Ozon-ecup-2026-user-ltv Public

    Прогноз GMV пользователей на 30 дней вперёд по истории поиска и каталога. Задача 3 хакатона E-CUP 2026 (Ozon), команда O3 — 14 место из 316. Ансамбль свёрточно-рекуррентных сетей и градиентного бус…

    Python 2

  2. Music-generation-fad-benchmark Music-generation-fad-benchmark Public

    Воспроизводимый бенчмарк 5 open-source text-to-music моделей по FAD, FAD-inf и CLAP Score на FMA-Pop и MTG-Jamendo. Три эмбеддера: CLAP, MERT, EnCodec. НИР МГТУ им. Н. Э. Баумана, каф. ИУ5.

    Python 1

  3. OreScope OreScope Public

    Классификация технологического сорта руды по OM-изображениям шлифов: сегментация фаз U-Net + интерпретируемое экспертное правило. macro-F1 0.91 по типу срастаний, панорама 300 Мп за 124 с. Хакатон …

    Python 1

  4. Smoking-detection-yolo-coreml Smoking-detection-yolo-coreml Public

    Детекция факта курения по видеопотоку: YOLO26n на классы person и smoke, факт фиксируется по пересечению bbox. mAP50 0.713, Precision 0.744. Экспорт в CoreML для инференса на iOS.

    Swift 1

  5. Stem-separator Stem-separator Public

    Разделение музыкального трека на 4 инструментальные дорожки на Demucs (htdemucs). REST API, Web UI, CLI, Docker. SDR 9.34 dB на вокале и 9.90 dB на басе (museval, MUSDB18), трек 3:17 за 14 с на M4 …

    Python 1

  6. X5-retailhero-uplift X5-retailhero-uplift Public

    Uplift на X5 RetailHero: кому предложение меняет поведение. Пересечение топ-10% по отклику и по приросту 0.44%. Δ uplift@10% = +9.53 п.п. [+6.63; +12.54].

    Python 1