Applied AI researcher

Hi :) I'm DhruviI teach small models to do big things
I explore how AI learns, reasons and acts, and turn that research into systems people can use. At JioHotstar, I build applied AI for real users. Outside work, research is my love.
Qwen2.5-3B · 1,000 held-out examples · tool-call accuracy
NMIMS · 2026Dean's Student ListCertificate of Merit
Top 10% of students
Say hi 👋
About
What happens if I try?
Underlined links open more
These days, most of my work comes down to two things:
A strict reward taught a 3B model to call tools, on one T4 GPU.
Read the paper →Ask a sports question and it writes the SQL, across Pro Kabaddi in five languages and ball-by-ball T20.
See the project →An HR coach that answers from handbooks, and a sales agent that picks shows for a brand and writes the deck.
HR coach →Sales agent →Research
4 AI research papers
Papers I once barely understood slowly turned into papers I wrote. Here's what each one is about, in plain words.
Advancing SLM Tool-Use Capability using Reinforcement Learning
- Cheap small models need exact tool-call formatting.
- GRPO training: any extra text earns zero reward.
- A 3B model reached 71.1% tool-call accuracy on one T4 GPU.
Towards Actionable Fashion AI: A Holistic, End-to-End Style Recommendation Assistant
- Seasonal colours and body shape become 20-25 clothing picks.
- Fine-tuned DINOv2: 61% seasonal-colour accuracy versus Gemini 2.5 Flash at 43%.
- Designed to run on a phone.
LightFusionRec: Lightweight Transformers-Based Cross-Domain Recommendation Model
- Recommend books from movies a user liked, even with little history.
- Combine DistilBERT text features with FastText genre features.
- MAE (mean absolute error): 0.69 at top 20%, beating text-only, genre-only and TF-IDF baselines.
Advanced AI-Based Detection and Tracking System (ADTS) for Crime Prevention and Identification in Real-Time Surveillance
- Detect people, recognise faces and track them across CCTV cameras.
- YOLOv8, MTCNN and FaceNet: 83 frames per second.
- IDF1 (identity-tracking score): 89.8%.
Awards & roles
Beyond the papers
College was never just about academics for me.
- Research Head, IEEE Robotics & Automation Society, NMIMS
- Student Coordinator, Student Research Group, NMIMS
- Sub-Head, Google Developer Student Clubs, NMIMS (400+ students)
Work
Things I've built
A mix of things running in production and experiments I just wanted answers to.
Text-to-SQL sports chatbots
- Started from zero, with a lot of research before building
- 12 seasons of Pro Kabaddi in English, Hindi, Hinglish, Gujarati or Marathi, plus ball-by-ball T20
- 96% right on our 150-question Kabaddi test
Cloud cost dashboard
- ~62 projects across GCP and AWS in one view
- About 4 days of manual work down to about 5 minutes
- Senior management uses it now
HR AI Coach
- Answers questions from company handbooks
- Keeps HR record facts separate from anything the AI writes, so the two never mix
Sales Agent
- Takes a brand's brief and picks the 5 best-fit shows
- Writes the pitch deck around them
- The sales team uses it now
Promo effectiveness dashboard
- Shows how a promo is actually landing
- Pulls signals from YouTube, Instagram, X, Reddit, Google Trends and Hotstar
- Reads the comments for sentiment
Episode condensation
- Works from the dubbing script
- Trims recaps, ad-break gaps and pauses
- One episode went from 24:33 to 13:56
Audio separation pipeline
- Benchmarked SAM-Audio against HTDemucs on a real production file; BS-RoFormer won out
- Splits TV audio into 6 stems, so crowd noise, applause and camera clicks come out while the music stays
- Ran on masters from 4 channels
Music catalog downloads
- Crawls a music library's full genre tree and saves each track as WAV to cloud storage
- Playwright automation, with ProcessPool running workers in parallel and systemd keeping long bulk downloads going
- A download lock so workers never clash, and every file checked against its track title and catalog ID
Music copyright detection
- Tested ACRCloud, AudioShake and Gemini against cue sheets I checked by hand
- AudioShake did best, 80.8-90.1% on music-only audio
- Moved timestamps out of the LLM and into Python, which took Gemini to 64%
HR onboarding videos
- Tried 5 lip-sync models, picked LatentSync
- With ElevenLabs voices and n8n, makes personalized onboarding videos in batches
Personal branding agent
- Co-built an LLM platform on n8n that writes, schedules and publishes social posts
- I designed the agent workflow nodes that make each post personal
Personalized news recommender
- Java and Spring Boot microservices with service discovery and an API gateway
- Runs on Docker and Kubernetes
- Recommendations from TF-IDF and each user's preferences
Multi-camera CCTV tracking
- Real-time face tracking across multiple CCTV feeds
- YOLOv8, MTCNN and FaceNet
- Became the ADTS paper
Side projects
Built for fun
- Hacker News → tweets. Every day it reads the top 15 Hacker News stories, sums them up and drafts 10 tweets from them. I built it with CrewAI and Streamlit. GitHub ↗
- Salary negotiation bot. It predicts your salary with ML, and then you haggle with an LLM bot over it. After that, a second bot negotiates the same profile with a boss bot, so you can see who got the better deal. GitHub ↗
Writing
Explained simply
Your model found the right answer. It might never learn to find it again.
What GRPO, a reward-based training method, changes across several tries.

The model weights are open. Can we see how the model learned?
Why open model training matters, and how to test the steps that shape an assistant.

A million tokens can still lose the right document
Why seeing every document is not the same as finding the right one.

Say hi