LatamBench Open Weights Observatory
The models Latin America can actually run.Evaluate open-weight and regional models under controlled serving. Measure accuracy, abstention, hallucination, quantization effects, serving variance, and benchmark-reference defects.
- Resource
- 300–600 A100-equivalent GPU-hours
- 90-day gate
- 6–8 open models across two benchmark families, with controlled repetitions and human validation.
- Public return
- Raw responses, exact revisions, serving recipes, judge transcripts, exclusions, cost, and measured utilization.
- Boundary
- Core scope stops at 70B. Larger models follow measured utilization, not ambition alone.