Skip to content
View TheDeadcoder's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report TheDeadcoder

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
TheDeadcoder/README.md

Nazmus Sakib

AI Engineer & Researcher

RL post-training · Agentic systems · Full-stack AI products

Portfolio Publication Hugging Face LinkedIn Codeforces Email


About

I build language-model systems end to end - the reinforcement learning that trains a policy, the agent architecture that puts it to work, and the product it ships inside.

My research is in RL post-training: reward design, GRPO and its variants, and the custom environments that make verifiable RL possible in domains where no benchmark exists yet. I'm drawn to the part usually skipped - building the simulator, the verifier and the reward signal, then measuring honestly whether RL earned its compute over supervised fine-tuning.

The engineering half is agentic. Memory architectures that let agents carry context across sessions (the subject of my WOA 2025 paper), tool calling and multi-step orchestration, graph and hybrid retrieval, and the adversarial testing that shows where those tool chains break. Around all of it sits the ordinary work that decides whether a model ever reaches a user: inference services, vector search, auth, payments, realtime web and mobile clients, and the CI that ships them. I also care about low-resource language evaluation, particularly Bengali.

CSE, BUET. Based in Dhaka, open to remote research collaboration.


What I Work On

Layer Focus
Research - RL post-training GRPO and variants (DAPO, Dr. GRPO, GSPO, RLOO) · multi-signal reward design · physics-grounded and verifiable environments · measuring RL headroom over SFT
Agentic systems Agent memory architectures · tool calling and multi-step orchestration · GraphRAG and hybrid retrieval · structured output and evaluation harnesses · adversarial testing of agent tool chains
Full-stack delivery Inference services and quantised local deployment · vector search · auth, payments and background jobs · Next.js / SvelteKit / Flutter clients · Docker and GitHub Actions
Reliability & evaluation Fault-injection scenarios for incident-resolution benchmarks · adversarial safety datasets for Bengali and other low-resource languages

Publication

MemAgent: A Cache-Inspired Framework for Augmenting Conversational Web Agents with Task-Specific Information N. Sakib, P. Barai, S. I. Parisa, A. Iqbal - WOA 2025, 26th Workshop From Objects to Agents, Trento, Italy. CEUR-WS Vol-4028, Paper 8

An agent memory architecture: a Memory Cache Bank with time-based expiration that decouples information gathering from task execution, so an agent stops re-asking users for details it has already learned. Reduces average conversation turns by 22.4% (5.00 → 3.88) across 150 Mind2Web tasks; a 15-participant study showed a 58% reduction in completion time for recurring tasks.


Selected Repositories

Reinforcement Learning & LLM Post-Training

Repository Description Stack
dc_ops_environment Physics-grounded datacenter RL environment - RC thermal networks and continuous multi-objective rewards, built on OpenEnv OpenEnv · Physics sim
dc_ops_training Teacher-distilled SFT → GRPO pipeline on AMD MI300X. Multi-signal reward (physics, scenario heuristics, anti-looping, format) drove +188% composite reward and a 10× per-step gain on hard multi-fault scenarios GRPO · TRL · Unsloth · vLLM · ROCm
GeoQL-4B Text-to-OverpassQL via SFT → GRPO on Qwen3-4B, with a documented negative result on reward redundancy in reference-based reward design GRPO · Qwen3 · vLLM
medical-cot-assistant Clinical chain-of-thought fine-tuning of a 20B model with QAT and LLM-as-a-Judge evaluation; INT4 + GGUF export for local inference Unsloth · QAT · llama.cpp

Agentic Systems & Retrieval

Repository Description Stack
civilmate-backend Agentic GraphRAG over building codes - plans multi-step lookups across a code graph, then compares design drawings against site imagery to generate technical logs and blocker alerts Neo4j · Qdrant · FastAPI
InsightAI-python-backend Microservices AI platform - tool-routed generation of quizzes and flashcards from PDFs and video, enforced structured output, and CLIP-based multimodal product search LlamaIndex · Qdrant · CLIP · FastAPI
Tokkhok-Backend Personalised RAG chatbot for "Banglish" (romanised Bengali) with a custom transliteration pipeline, few-shot inference and configurable agent personas FastAPI · Qdrant · PostgreSQL
sust-backend Bebsha AI service layer - RAG product search, description generation, background removal Flask · RAG

Applied Deep Learning

Repository Description Stack
bd-prescription-medicine-recognize ResNet50–CRNN for 78-class handwritten medicine-name recognition on Bangladeshi prescriptions. Hash-grouped StratifiedGroupKFold + SWA ensemble reached 92.53% test accuracy / 0.9234 macro-F1 (+6.44 pt over baseline) PyTorch · CRNN · SWA · MLflow · Grad-CAM

Systems & Reliability

Repository Description Stack
SREGym Contributed fault-injection scenarios to an incident-resolution benchmark for AI agents - Kafka poison-pill head-of-line blocking, CFS throttling brownout, and oscillating config corruption Kubernetes · Kafka · Chaos engineering

Full-Stack Products

Repository Description Stack
nerdherd2ndrun Collaborative productivity platform - shared notes, video calls, real-time quizzes, AI assistant SvelteKit · Firebase
coderhub Developer community platform with blogging, skill-based search, and project management SvelteKit · Vercel
yobofrontend Frontend for YoboSQL, a text-to-SQL conversational interface TypeScript · SvelteKit

Engineering Templates

Repository Description Stack
django-init-template Production-ready Django REST scaffold with PostgreSQL and Supabase auth Django · DRF · Supabase
nodejs-init Express + TypeScript service scaffold with Helmet, Swagger, and Supabase auth TypeScript · Express

Open Models & Datasets

Published on Hugging Face @Melikshah.

Artifact Description
GPT-OSS-20B-Clinical-CoT (GGUF) 4-bit quantised 20B model fine-tuned for clinical chain-of-thought reasoning, optimised for local inference
dc_ops_grpo_lora GRPO-trained LoRA adapter for the datacenter operations environment
qwen3.5-4b-base-blindspots Adversarial evaluation set probing architectural and logical failure modes of Qwen3.5-4B-Base
Shajgoj & General Products 20,000+ item multimodal datasets for image–text retrieval
Prothom Alo News Large-scale Bengali news corpus for low-resource NLP research

Experience

Role Organisation Focus
Lecturer, CSE BRAC University · present Machine Learning, Operating Systems, Software Engineering, System Analysis & Design
Founding Engineer (Backend & AI) Intellesphere · 2024–25 Banking RAG with hybrid indexing and Keycloak-secured access, CLIP multimodal search, multimodal civil-engineering automation
Founding Engineer (Backend & AI) Oleyn · 2024 Bengali ASR with NeMo speaker diarization, agentic CRM and campaign system, legal research assistant
Software Engineer (Backend) Priyo · 2024 Django support agent at concurrency, warehousing REST APIs, campaign and analytics infrastructure
Software Engineer Intern Yobo · 2024 Text-to-SQL chat interface on LangChain, FastAPI and SvelteKit

Technical Stack

Domain Tools
Languages Python · C++ · Java · TypeScript · JavaScript · SQL
Training & Inference PyTorch · TRL · Unsloth · PEFT/LoRA · vLLM · llama.cpp · ROCm · Transformers · NeMo · scikit-learn · MLflow
Reinforcement Learning GRPO family (DAPO, Dr. GRPO, GSPO, RLOO) · reward design · custom verifiable environments · OpenEnv
Agentic Systems Tool calling · agent memory · multi-step orchestration · LangChain · LlamaIndex · Genkit · RAG · GraphRAG · hybrid retrieval · structured output · eval harnesses
Backend & APIs FastAPI · Django · Spring Boot · Express · Firebase Cloud Functions · Nginx · Keycloak · Stripe
Frontend & Mobile Next.js · React · SvelteKit · Flutter · Tailwind CSS
Data & Vector Stores PostgreSQL · MySQL · Redis · Qdrant · ChromaDB · Neo4j · Supabase · Turso
Infrastructure Docker · Kubernetes · AWS · GCP · Vercel · Modal · GitHub Actions · PostHog

GitHub Activity

GitHub activity overview Language distribution


Selected Honours

Award Event Year
Champion IUT 11th National ICT Fest - OpenAPI Hackathon 2024
Champion CodeCrafters Dev Sprint Hackathon, BUET 2024
Champion SUST CSE Carnival Hackathon 2024
Runner-Up Gen-Dev Hackathon, Acme AI 2024
Honourable Mention Bangladesh Blockchain Olympiad (BCOLBD) · IDSOL World Finalist 2024
Champion Cefalo ITverse Project Showcase 2023

Projects · Models · Get in touch

Pinned Loading

  1. dc_ops_training dc_ops_training Public

    Teaching a 7B Model to Run a Datacenter

    Python

  2. medical-cot-assistant medical-cot-assistant Public

    Supervised fine-tuning pipeline for medical reasoning using SFT, QAT, and LLM-as-a-Judge evaluation

    Python

  3. Tokkhok-Backend Tokkhok-Backend Public

    Demo (https://www.youtube.com/watch?v=2UInbtVz1oE)

    Python 5 1

  4. InsightAI-python-backend InsightAI-python-backend Public

    This is the Python backend for InsightAI, Architected a microservices-based EdTech platform combining Study Companion, Project Planner, and ShopGenie modules.

    Jupyter Notebook 4

  5. civilmate-backend civilmate-backend Public

    Engineered an agentic retrieval system combining Neo4j (Graph) and Qdrant (Vector) to interpret complex building codes. Automated site reporting by processing construction images with multimodal AI…

    Python

  6. IUTCS-24-OpenAPI-Hackathon/BUET_Genesis IUTCS-24-OpenAPI-Hackathon/BUET_Genesis Public

    This project focuses on solving some challenges using third party free/freemium APIs. Here we have used APIs like GeoApify, WeatherAPI, Amadeus API etc. Using these APIs we implemented geolocation,…

    Svelte 1