# ScaleDown > ScaleDown builds task-specific small language models (SLMs) for compression, summarization, extraction, and classification. 10x lower cost, 10x lower latency than general-purpose models. ## What is ScaleDown? ScaleDown (scaledown.ai) is an applied AI research lab building purpose-built small language models for specific tasks. Each model is trained to do one thing extremely well, at a fraction of the cost and latency of general-purpose alternatives. ScaleDown models deliver frontier-quality results while being 10x cheaper and 2x faster than models like GPT-4 or Claude. ## Models ScaleDown offers four task-specific models via a unified REST API: ### COMPRESS Lossless, query-aware compression. Strip the noise out of long context without losing the signal. Send in a long prompt or document. Get back a semantically compressed version that preserves every fact relevant to your query, at a fraction of the token count. Typical compression ratio: 50-70% Real impact: A workflow costing $10/day drops to $4/day at 60% compression Use cases: - RAG pipelines: compress retrieved chunks before passing to your LLM - Document analysis: process full documents without truncation - Conversation management: condense history for long sessions - Code review: fit large codebase context into a single call - Batch workflows: reduce costs at scale - Reasoning models: cut overhead on chain-of-thought traces ### SUMMARIZE Abstractive summarization at frontier model quality, 10x lower price, 10x lower latency. Built for pipelines that need human-quality summaries at scale. Handles long documents, transcripts, and threads without truncation. Use cases: - Legal contracts and research papers - Customer support tickets and surveys at scale - Meeting transcripts and action item extraction - Earnings calls and financial filings - Content pipelines: prepare long articles for downstream processing ### EXTRACT Semantic NER using natural language. Define what you want in plain English and pull it out as structured data. Describe your entities: names, dates, product SKUs, any schema you define. Returns clean JSON every time. Use cases: - Contract analysis: parties, dates, obligations, clauses - Resume parsing: skills, employment history, education - Medical record processing: diagnoses, medications, procedures - Financial document review: figures, entities, references - E-commerce: products, prices, SKUs from unstructured text ### CLASSIFY Sort data fast. Built for high-throughput tagging, triage, and filtering. Define your label taxonomy in plain text. Classify millions of items per hour with consistent, calibrated outputs. No fine-tuning required. Use cases: - Content routing: direct support tickets to the right team - Moderation: confidence-scored safety layers with custom thresholds - Intent recognition: questions, complaints, cancellations, feedback - Domain tagging: classify documents by topic without training data - Lead evaluation: score inquiries by urgency or deal fit ## Benchmarks ScaleDown's task-specific models are benchmarked against frontier general-purpose models from OpenAI, Anthropic, Google, and xAI on public academic benchmarks (FinanceBench, QMSum, CUAD, CLINC OOS). Metrics: F1 for FinanceBench/CUAD, ROUGE-L for QMSum, accuracy for CLINC OOS. Cost is per 1,000 calls at list pricing; latency is mean response time per call. Benchmark index: https://scaledown.ai/benchmarks ### ScaleDown vs OpenAI (GPT-5.5, GPT-5.4 Mini, GPT-5.4 Nano) https://scaledown.ai/benchmarks/scaledown-vs-openai Average across 4 tasks: 8.72% higher accuracy, 89x cheaper, 2.4x faster than the OpenAI average. - Compress (FinanceBench): ScaleDown 0.311 F1 vs OpenAI avg 0.249 (+6.23%), 138x cheaper, 1.85x faster. Closest race on the page: GPT-5.5 scores 0.307, within half a point. - Summarize (QMSum): ScaleDown 0.267 ROUGE-L vs OpenAI avg 0.179 (+8.83%), 41x cheaper, 2.67x faster. - Extract (CUAD): ScaleDown 0.612 F1 vs OpenAI avg 0.468 (+14.40%), 35x cheaper, 3.26x faster. Widest accuracy gap in the OpenAI comparison. - Classify (CLINC OOS): ScaleDown 0.950 accuracy vs OpenAI avg 0.896 (+5.40%), 1,810x cheaper, 2.57x faster. ScaleDown beats every individual GPT model, not just the average. ### ScaleDown vs Anthropic (Claude Fable 5, Opus 4.8, Sonnet 4.6, Haiku 4.5) https://scaledown.ai/benchmarks/scaledown-vs-anthropic Average across 3 tasks: 8.0% higher accuracy, 161x cheaper, 3.8x faster than the Anthropic average. - Summarize (QMSum): ScaleDown 0.267 ROUGE-L vs Anthropic avg 0.180 (+8.22%), 217x cheaper, 4.4x faster. - Extract (CUAD): ScaleDown 0.612 F1 vs Anthropic avg 0.480 (+13.33%), 91x cheaper, 3.2x faster. Widest accuracy gap in the Anthropic comparison. - Classify (CLINC OOS): ScaleDown 0.950 accuracy vs Anthropic avg 0.930 (+2.29%), 5,250x cheaper, 3.6x faster than the average. Opus 4.8 and Fable 5 edge slightly higher on raw accuracy alone (0.989 and 0.979), but at thousands of times the cost and several times the latency for a marginal gain. ### ScaleDown vs Gemini (Gemini 3.1 Pro, 3.5 Flash, 3.1 Flash-Lite) https://scaledown.ai/benchmarks/scaledown-vs-gemini Average across 3 tasks: 9.0% higher accuracy, 29x cheaper, 8.3x faster than the Gemini average. - Summarize (QMSum): ScaleDown 0.267 ROUGE-L vs Gemini avg 0.120 (+14.70%), 28x cheaper, 2.6x faster. Widest accuracy gap on this page. - Extract (CUAD): ScaleDown 0.612 F1 vs Gemini avg 0.509 (+10.34%), 26x cheaper, 13.4x faster. - Classify (CLINC OOS): ScaleDown 0.950 accuracy vs Gemini avg 0.929 (+2.10%), 2,437x cheaper, 13.8x faster. Closest race on this page: Gemini 3.1 Pro scores 0.949, within a tenth of a point. ### ScaleDown vs Grok (Grok 4.3, Grok 4.2) https://scaledown.ai/benchmarks/scaledown-vs-grok Average across 3 tasks: 7.93% higher accuracy, 24x cheaper, 2.2x faster than the Grok average. - Summarize (QMSum): ScaleDown 0.267 ROUGE-L vs Grok avg 0.167 (+10.05%), 25x cheaper, 2.1x faster. - Extract (CUAD): ScaleDown 0.612 F1 vs Grok avg 0.489 (+12.30%), 22x cheaper, 2.3x faster than the average. Grok 4.2 is marginally faster on this task alone (861ms vs 918ms), but trails ScaleDown by double digits on accuracy at roughly 21x the cost. - Classify (CLINC OOS): ScaleDown 0.950 accuracy vs Grok avg 0.936 (+1.45%), 1,161x cheaper, 2.3x faster. Closest accuracy race on this page. ## API Base URL: https://api.scaledown.xyz Endpoints: - POST /compress - POST /summarize - POST /extract - POST /classify Authentication: x-api-key header Documentation: https://docs.scaledown.ai ## Pricing Public API: $0.05 per 1M tokens - 50M free tokens to start, no credit card needed - All four models included - Flat pricing, no tiers - Generous fair-use rate limits Enterprise / Self-hosted: custom pricing - Self-hosted in your VPC - Fine-tuning on your data - Dedicated support and SLAs ## Compliance and Security - Zero data retention: prompts, outputs, and API keys are never logged or stored. In progress with GDPR, HIPAA, and SOC 2 compliance. - No model training on your data: your inputs are never used to fine-tune or improve ScaleDown models. - Self-hosted deployment: run entirely in your VPC. No outbound calls, no shared infrastructure, full air-gap support. - Enterprise access controls: SSO, RBAC, and audit logs available on enterprise plans. ## Infrastructure Partners Amazon Web Services (AWS), Google Cloud, NVIDIA, Intel, Nutanix ## Blog Technical posts on SLMs, benchmarking methodology, and applied AI research. Index: https://scaledown.ai/blog - How We Benchmark Extract on CUAD — https://scaledown.ai/blog/benchmarking-scaledowns-extraction Extract vs. GPT-5.4 Mini on 510 real commercial contracts (CUAD): +7.65 EM, +11.43 F1, 60% cheaper. - How We Benchmark Summarize on QMSum — https://scaledown.ai/blog/benchmarking-scaledowns-summarization Summarize vs. frontier models on meeting-transcript QA (QMSum): comparable ROUGE-L, 4-51x cheaper. - Benchmarking Compress on FinanceBench — https://scaledown.ai/blog/benchmarking-context-compression Compress vs. a GPT-5.2 baseline on FinanceBench via ScaleBench: +6.5% F1 and 61% lower cost at 70% compression. - Benchmarking Classify for Model Routing — https://scaledown.ai/blog/benchmarking-scaledowns-classification Classify vs. GPT-5.4 Mini/Nano on intent classification for LLM routing: 90.53% accuracy at $0.0008, beating both baselines on accuracy and cost. - Zero Data Retention at ScaleDown — https://scaledown.ai/blog/zero-data-retention-at-scaledown What ZDR means in practice: prompts are tokenized in memory, processed, and discarded — never written to disk, in any environment, including on failure. - No Model Training on Your Data at ScaleDown — https://scaledown.ai/blog/no-model-training-on-your-data-at Customer inputs are never used to train, fine-tune, or improve ScaleDown models, during or after inference, in any form. - Self-Hosted Deployment of ScaleDown — https://scaledown.ai/blog/self-hosted-deployment-of-scaledown When and how to run ScaleDown entirely inside your own VPC for data that can't leave your environment. - Task-Specific Models are the New Frontier — https://scaledown.ai/blog/task-specific-models-are-the-new Why purpose-built small models are replacing the "big general model + better prompts + RAG" default enterprise AI strategy. - How We Train Small Language Models for Context Compression — https://scaledown.ai/blog/how-we-train-small-language-models How the Compress endpoint's task-specific SLM is trained to identify which parts of a context are relevant to a given query. - Intent Classification for AI Financial Planners — https://scaledown.ai/blog/intent-classification-for-ai-financial Why a task-specific classifier beats a frontier LLM for routing financial-planning queries on cost, latency, and domain over-interpretation. - Context Compression for Data Security and Privacy Platforms — https://scaledown.ai/blog/context-compression-for-data-security Applying Compress to platforms that trace sensitive data across source code, infrastructure, and AI models. - Extraction SLMs for E-Commerce Chatbots and Voice Agents — https://scaledown.ai/blog/extraction-slms-for-e-commerce-chatbots Why structured-data extraction for e-commerce support doesn't need frontier LLM pricing. ## Resources - Website: https://scaledown.ai - Docs: https://docs.scaledown.ai - Get API key: https://scaledown.ai/auth - Blog: https://scaledown.ai/blog - Contact: contact@scaledown.ai