Skip to content

Moinuddin Shaik · AI Engineer · Applied Scientist · Hyderabad

I build AI systems, and the evaluation that proves they work.

Applied Scientist Intern at Amazon RBS Sciences. First author on a submitted paper, and nine systems written up with the numbers attached.

Selected work

Systems where correctness had to survive contact with reality.

Production LLM work at Amazon, a grading-alignment eval against the human ceiling, first-author taxonomy research, and citation-first retrieval. Each one states its constraint before its stack.

Applied science · Public summary2026

Autonomous Taxonomy Systems at Amazon

A self-calibrating applied AI system that turns large-scale customer feedback into explainable three-level taxonomies—without a scientist hand-tuning every new domain.

0.74

F1 vs. 0.71 baseline

<18h

onboarding cycle

2.7×

faster extraction

Results

Measured against a baseline, every time.

A number without a comparison is a claim. Each row is what I built against what it replaced, measured the same way — and links to the write-up with the method.

From Amazon

He worked on using LLMs for taxonomy use cases, he is a remarkably quick learner who brings new ideas and executes them fast.
Manan Soni, Applied Scientist II at Amazon · Mentored Moin during the internship

From Amazon

His passion for solving complex problems stood out from day one. He took on a genuinely challenging project and delivered real impact, backing every decision with thoughtful, well-run experiments.
Sachin Giroh, Applied Scientist at Amazon · Worked with Moin on the same team at Amazon

Experience

A short record of outcomes, not job descriptions.

What was broken, what I built, the decisions that mattered, and what measurably changed.

The problem

Creating a taxonomy for a new feedback domain took a seven-notebook workflow. Scientists picked examples, tuned prompts, ran clustering, and stitched a three-level hierarchy by hand — five to seven days of expert time per domain, with a scientist in every step.

What I built

A self-calibrating knowledge-extraction and taxonomy-induction system, shipped as a production container for multi-domain, million-record workloads.

Approach

  • Mine and validate a small, diverse in-domain example set, then run cached batch extraction at scale.
  • Density-based L3 discovery over structured phrases; the LLM classifies those clusters into a disjoint hierarchy.
  • Deterministic attribution computes every quality claim before the LLM renders it in plain language.

What changed

0.74
taxonomy F1, against a 0.71 manual scientist baseline
<18h
autonomous run, down from a five-to-seven-day workflow
23.8% → 0.7%
missing Category and Aspect extractions
2.7× / 53%
faster extraction, lower inference cost
12 teams
adopted the explainable evaluation framework
Moinuddin Shaik at the Amazon office in Bengaluru, wearing a visitor badge, standing behind the lobby's I-heart-amazon lettering

1 of ~200 Applied Scientist interns across India. Selected via Amazon ML Summer School 2025 — 3,000 of 1.6 lakh applicants.

Read the full case study

The problem

Real-time video enhancement had to run on low-resource devices with no dedicated GPU, where the usual answer is to accept worse output or a slower frame budget.

What I built

An optimized real-time video enhancement pipeline targeted at CPU-only inference.

Approach

  • Traded model capacity against perceptual quality rather than accepting the default architecture.
  • Tuned for CPU inference as the target, not as a fallback path.

What changed

20%
clearer output
30%
smaller model
35%
faster CPU inference, without a dedicated GPU

What I build

Four kinds of problem, and the work behind each one.

Grouped by the problem, not the technology. Pick one and the evidence changes beside it.

Built with — every tool below is in at least one shipped project; ×n is how many

ML & modellingTraining, fine-tuning, and evaluation
  • PyTorch
  • Transformers
  • TensorFlow Lite
  • ONNX
  • Whisper-tiny
  • BERTopic
  • UMAP
  • HDBSCAN
MethodsHow the work is done, not what it is done with
  • Experiment design×2
  • Classification
  • Hierarchical clustering
  • NLP
  • LLM evaluation×2
  • Taxonomy evaluation
  • Knowledge extraction
LLM systemsRetrieval, agents, and the evaluation around them
  • LangGraph
  • pgvector×2
  • Docling
  • LLM systems
Backend & dataThe services the models actually run inside
  • Python×6
  • FastAPI×5
  • PostgreSQL×4
  • Redis×3
  • Celery
  • Temporal
  • Docker
  • Pydantic
  • AWS
  • Appwrite
InterfacesWhere the output has to be legible to a person
  • React
  • React 19
  • React Native
  • Expo
  • Remotion

Lab

How Ask finds an answer without a model.

Not a simulation. Your question is tokenised in the browser, each term weighted by how rare it is across this site's own text, and every document ranked by cosine similarity. Nothing hidden: type, and watch each number change.

tf-idf · cosine · 16 docs · 1009 terms

runs in your browser · no model call

w(t) = tf(t) · (ln((N+1)/(df(t)+1)) + 1)·score(d) = Σt q̂(t) · d̂(t)·q̂, d̂ unit length, so the sum is the cosine

Publications

Research that makes model behavior easier to measure.

Current work focuses on hierarchy quality, classification at large label scales, and the cost of reliable decisions. Submission status is stated plainly.

Amazon ML Conference · 2026Submitted

Universal Anecdote Miner (UAM)

First author

A framework for evaluating how automated systems construct hierarchical structure from raw feedback, including analysis of a subtle duplication failure mode in human-built taxonomies.

EMNLP · 2026Submitted

LUMEN: Robust LLM Classification Across Taxonomy Scales

Third author

A scale-aware classification study spanning 150 to 5,000 labels. I contributed the taxonomy-scaling experiments; on Clothing, LUMEN maintained Claude-competitive F1 at up to 99% lower inference cost than Sonnet 4.5.

Open source

The implementation is part of the argument.

Public repositories include product code, typed APIs, tests, CI, deployment notes, and explicit limitations—not only screenshots.

Before AI

Two things I got good at before this one.

Both are real and both are measured by someone other than me. Open either if you want the detail.

Cinematic editing and motion graphics — Fiverr first, then direct, over roughly five years alongside school and college. 4.9 across 117 reviews. Nobody assigned it: he found the work, taught himself the craft, and delivered to a brief on a deadline for people who were paying.

Fiverr gig listing for cinematic video editing, rated 4.9 across 117 reviews
The account is my sister's — I was too young to hold one. Her name and photo are redacted at her request.

Picked up badminton at 11, competed at state level, took a doubles silver at Student Nationals in 2018. Lockdown ended it — but it was the first time getting good at something was measured by someone other than him.

The pattern

Start early. Go deep. Repeat.

Different domains, different kinds of proof: competitive, commercial, published.

11

Badminton

State-level competition. Student Nationals silver in doubles, 2018. Stopped during lockdown.

14

Code

Started in 10th class, 2020.

15

Video

Editing for clients on Fiverr, then direct. 200+ clients over roughly five years.

20

Amazon

Applied Scientist Intern, 1 of ~200 across India. Selected via Amazon ML Summer School 2025 — 3,000 of 1.6 lakh applicants.

Now

Now

B.Tech CS (AI & ML), 9.09/10. Building AI systems, products, and research.

Before AI

Started editing at 15. Built for 200+ clients.

Cinematic video editing and motion graphics for clients worldwide. Fiverr first, then direct — across roughly five years, alongside school and then college.

4.9

Rating across 117 Fiverr reviews

200+

Clients over roughly five years, Fiverr then direct

15

Age I started, working under my sister's account

I was 15, which is too young to hold a seller account — it needs legal documents I did not have — so I worked under my sister's. Fiverr first, then direct clients as the work grew. None of it was assigned to me: I found the work, taught myself the craft, and delivered to a brief on a deadline for people who were paying. Scope, revisions, and clients across timezones taught me the parts of building that have nothing to do with code, years before I had a job title. The gig is still up and I take the occasional project, but AI is the work now.

Fiverr gig listing for cinematic video editing, rated 4.9 across 117 reviews
The gig as it stands today. The account is my sister's — her name and photo are redacted here at her request.

Six of the 117.

Fiverr review, five stars, United States: a client names Moinuddin and cites audio editing and 3D motion graphics
Fiverr review, five stars, United States: praises creative solutions and a 24-hour turnaround other editors declined
Fiverr review, five stars, Australia: repeat client says they will use the service again
Fiverr review, five stars, United Kingdom: calls the seller a consummate professional
Fiverr review, five stars, Mexico: praises excellent work and punctual delivery
Fiverr review, 4.3 stars, United States: notes talent and resourcefulness alongside criticism of communication

United States, Australia, United Kingdom, Mexico. Including a 4.3 — the average is 4.9.

Before that

Student Nationals. Silver, doubles.

Silver.

I picked up badminton at 11 and competed at state level before taking a doubles silver at Student Nationals in 2018. Lockdown ended it, but it was the first thing I got properly obsessed with — and the first time getting good at something was measured by someone other than me.

11
Age I started
2018
Nationals silver
State
Level competed at