Skip to content
Selected work
Research · UAM + LUMEN · 2026

Evaluation for Taxonomies at Scale

As label spaces grow, evaluation must distinguish genuine structure from duplicated concepts, brittle boundaries, and expensive model behavior.

First author, UAM · Third author, LUMEN

621–5K

category scale studied

Up to 99%

lower cost than Sonnet 4.5

2

2026 submissions

01

Problem

A taxonomy can look coherent while repeating the same idea across branches. At the same time, classification systems that work at small label counts can become costly or unstable as the taxonomy grows.

02

Why it matters

Teams use taxonomies to aggregate evidence and decide what to fix. Structural duplication distorts counts, while classification cost can make an otherwise accurate system impractical.

03

Architecture

6 stages · select to inspect

UAM focuses on evaluating generated hierarchical structure and surfacing duplication. My LUMEN contribution focused on scaling taxonomies up and down and evaluating how classification quality, hallucination, and cost change as the label space grows.
04

Technical challenges

01

Measuring hierarchy, not just labels

Designed analysis around parent-child structure, overlap, and duplicated concepts rather than only flat accuracy.

02

Holding comparisons fair

Contributed taxonomy-scaling experiments from 621 to 5,000 Clothing categories while keeping the input domain fixed, then compared quality and inference cost under increasing label-space complexity.

03

Turning anomalies into findings

Connected quantitative structure checks to a failure mode people could inspect and reason about.

05

Tradeoffs

Interpretable measures over one composite score

Separate signals made structural failure modes visible and actionable.

Scale sweep over a single benchmark point

A model that is practical at hundreds of classes may behave differently at thousands.

Submission status stated explicitly

The portfolio does not imply acceptance or publish details that are still under review.

06

Experiments

  1. 01Compared hierarchy evaluation behavior across generated and human-constructed structures.
  2. 02Contributed taxonomy-size sweeps that expanded the Clothing label space from 621 to 5,000 categories.
  3. 03Compared LUMEN with Claude Sonnet 4.5 and Qwen baselines while tracking F1, hallucination, and inference cost.
07

Results

UAM identified a previously hidden duplication failure mode in human-built structures.

On Clothing, LUMEN reached 90.3 F1 versus Sonnet 4.5's 90.8 at the original taxonomy, with roughly 96% lower inference cost.

At 5,000 categories, LUMEN reached 89.6 F1 versus Sonnet 4.5's 89.4, with roughly 99% lower inference cost.

UAM is an AMLC 2026 submission with me as first author; LUMEN is an EMNLP 2026 submission with me as third author.

08

Lessons learned

  • Human-authored structure is a baseline, not ground truth beyond inspection.
  • Cost belongs in the scientific result when it changes deployability.
  • Error analysis is most valuable when it reveals a repeatable failure class.
09

Future work

  • Release public artifacts when review and confidentiality constraints allow.
  • Extend structural measures to deeper and evolving hierarchies.
  • Study how taxonomy quality changes downstream prioritization decisions.