EXCEEDS logo
Exceeds
Xiang An

PROFILE

Xiang An

Developed and integrated the VisFactor benchmark task into the EvolvingLMMs-Lab/lmms-eval framework, delivering a rule-based scoring pipeline that normalizes multiple answer types and mirrors the VLMEvalKit standard. The work included implementing two-level aggregation for evaluation parity on thousands of real-model predictions, as well as creating new task configurations and data-processing utilities. Leveraging Python, YAML, and regex, the solution enables reproducible, vendor-agnostic benchmarking of vision-cognition models while reducing reliance on external LLM judges. Comprehensive documentation and parquet data conversion references were provided, streamlining data ingestion and supporting quick adoption within the open-source evaluation ecosystem.

Overall Statistics

Feature vs Bugs

100%Features

Repository Contributions

1Total
Bugs
0
Commits
1
Features
1
Lines of code
284
Activity Months1

Work History

June 2026

1 Commits • 1 Features

Jun 1, 2026

June 2026 delivered the VisFactor Benchmark task integration into the EvolvingLMMs-Lab lmms-eval framework. The feature adds a rule-based VisFactor scoring pipeline with normalization for multiple answer types, a new task configuration, supporting data-processing utilities, and comprehensive documentation. This integration aligns evaluation with the VLMEvalKit standard, enabling reproducible, vendor-agnostic comparisons of vision-cognition models and reducing dependency on external LLM judges. Work includes validation against real model predictions and sets up data wiring for parquet-based assets stored in the HF Hub reference. Repo: EvolvingLMMs-Lab/lmms-eval.

Activity

Loading activity data...

Quality Metrics

Correctness100.0%
Maintainability100.0%
Architecture100.0%
Performance80.0%
AI Usage60.0%

Skills & Technologies

Programming Languages

No languages yet

Technical Skills

BenchmarkingData ProcessingPythonRegexYAML

Repositories Contributed To

1 repo

Overview of all repositories you've contributed to across your timeline

EvolvingLMMs-Lab/lmms-eval

Jun 2026 Jun 2026
1 Month active

Languages Used

No languages

Technical Skills

BenchmarkingData ProcessingPythonRegexYAML