← All Materials
← → arrow keys
THAGORUS PBC
$1.25M Pre-Seed

Your brand has 10 years of data.
What if you could train on
10,000?

AI transformed language, biology, and self-driving by pooling data at scale.
We're building a shared model trained across thousands of brands —
so every brand gets smarter together.

+16% accuracy improvement from cross-brand training — already proven
thagorus.com
The Opportunity

ML transformed every major industry.
One conspicuous gap remains.

ML BREAKTHROUGHS LLMs / Agents Autonomous agents plan and execute multi-step tasks — Gemini 3, Claude 4 Drug Discovery 173+ AI-designed drugs in clinical trials. 80-90% early success vs 52% historical Autonomous Vehicles Waymo targeting 1M rides/wk across 27 cities Healthcare / Clinical 90% of hospitals adopting AI diagnostics Creative / Video Sora 2, Kling 3.0 — native audio, real controls Software Eng. $12.8B market. Cursor, Claude Code, Copilot X Climate & Weather NOAA deployed AI forecasts. 99.7% less compute Financial Markets Multi-agent trading systems, on-prem SLMs Economic Superintelligence $5.5T in US retail. Every brand is an island. Still guessing alone. ? $5.5T in retail. No one has pooled the data.
The Problem

450,000 brands sitting on demand data.
Each one guessing alone.
None of them can see the full picture.

NorSari
NorSari
Blanket sales spike 40% in a cold snap. They think it's their Instagram ad. They double the ad budget. Weather changes. Sales drop. They never learn why.
Homesick Candles
Homesick Candles
Demand surges every October. Is it the campaign? The gifting season? The first frost? They can't tell. So they overstock, overspend, and repeat. Every year.
×
450,000 BRANDS
Every other brand
Same story, different product. Each sitting on years of demand data. But no brand alone has enough data to see the full picture. Together, they would.
GPT works because it trained on all the text, not just one book. Retail demand data has the same property — but nobody has pooled it yet.
$1.8T
annual demand misjudgment
across US retail
Why Now

Scaling laws reached
economic data
this year.

Every major AI breakthrough followed the same pattern: pool more data, train a bigger model, get a step-function improvement. This year, scaling laws were proven for economic time series — meaning the same approach that built GPT now applies to demand prediction.

2017 2019 2020 2022 2023 2024 2026 TIME → CAPABILITY → GPT-2 GPT-3 GPT-4 Claude 4 / Gemini 3 Language / Agents Autonomous multi-step reasoning DALL-E 2 Midjourney Sora 2 / Kling 3.0 Video / Creative Native audio, real cinematography controls AlphaFold AF3 Zasocitinib Ph III Drug Discovery 173+ AI drugs in clinical development Waymo Waymo 27 cities Autonomous 1M rides/wk target, Zoox launched Cursor Claude Code Code $12.8B market, full app generation WeatherVane ECONOMIC SUPERINTELLIGENCE Every discipline follows the same curve. More data → bigger model → step-function breakthrough
SCALING LAWS
Proven for time series
NeurIPS 2024. Same physics as LLMs now applies to demand.
DATA ACCESS
Commerce APIs opened
Shopify, Amazon expose structured demand data via API.
FREE SIGNALS
Context data is free
Weather, macro, competitive signals at zero cost via open APIs.
BROKEN ATTRIBUTION
Brands need a new source
iOS 14.5 killed attribution. Brands need intelligence from outside their silo.
The Proof

Cross-source training works
for time series. We've already
proven it on retail.

Time-series foundation models are moving fast — Google, Amazon, and Salesforce all published major papers in 2024-25. The key finding: training on data from many sources makes predictions for each source dramatically better. We applied this to retail demand and it works.

DEMAND PREDICTION ACCURACY
Single brand alone
Brands pooled together
+16%
Tested on M5 — the world's largest retail forecasting benchmark (100K Walmart time series). Our improvement is 4x what's considered significant in the research community.
RESEARCH BREAKTHROUGHS (2024–2026)
Google TimesFM — proved scaling laws transfer to time-series forecasting. NeurIPS 2024.
Amazon Chronos — zero-shot forecasting works across domains without fine-tuning.
Salesforce Moirai — multi-domain transfer across 27B observations. Cross-source training validated.
Nixtla TimeGPT — first commercial time-series foundation model. Proved market demand exists.
WEATHERVANE — BUILT & PROVEN
✓ +16% accuracy from cross-brand pooling — 4x the benchmark significance threshold
✓ Causal attribution engine — isolates weather, promos, competition, and baseline for each brand
✓ Two design partners with 8+ years of demand data (NorSari, Homesick Candles)
✓ Full pipeline built — ingestion, training, prediction, serving. Solo technical founder.
What this means in dollars: For a $10M brand, +16% forecast accuracy is the difference between overstocking by 12% vs. 3%. That's hundreds of thousands in recovered margin — per brand, per year.
THIS HAS BEEN DONE BEFORE:
IMS Health → IQVIA Pooled pharmacy data across rivals. Now worth $50B.
Renaissance Technologies Pooled diverse data (weather, satellite, economic). 66% avg annual returns.
Bloomberg Aggregated financial data no firm would share. $70B private company.
Verisk Pooled 32B insurance records across competitors. $35B market cap.
Nielsen Built retail measurement by pooling scanner data across brands.
The Flywheel

Every brand
makes every
other brand
smarter.

When a brand joins, its data improves predictions for every other brand. And it gets accurate predictions from day one, before contributing anything — because the corpus is already working.

Data network effect
More brands = better model = attracts more brands. +16% accuracy already proven from cross-domain training.
Inverted cold-start
New brands get predictions powered by all existing data before contributing anything. They see the value on day one.
Historical corpus lock-in
Early movers own the corpus permanently. You cannot buy 10 years of winter demand data in 2028.
A competitor starting in 2027 cannot go back and buy 10 years of everyone's winter demand patterns. That data either exists in the corpus or it doesn't.
MORE DATA BETTER MODEL BETTER PREDICTIONS MORE BRANDS
+16%
accuracy from
cross-domain
4x
M5 significance
threshold
10yr
head start on
historical corpus
The Product

What the network
unlocks for each brand

These outputs only work because of the cross-brand corpus. For every SKU x location x week:

Powered by the cross-brand network: Every output below gets more accurate as more brands join. Each brand's data improves the model for all brands.
1
Real-time timing signal
"Pause ads this week" — your market is about to cool. Act-now / wait recommendations that save ad budget.
2
SKU-level demand forecast powered by the network
Point forecast + uncertainty interval, refreshed weekly. Trained on the pooled corpus, not just your data — so it's accurate from day one.
3
Causal attribution (the "why")
How much was weather vs. promo vs. competition vs. baseline? Patterns only visible when you can compare across hundreds of brands.
WeatherVane · NorSari · Cozy Knit Collection · Northeast US · Wk of Mar 24
Demand Forecast
Oct Nov Dec Now Feb
← 6 weeks back this week ↑ 3-week fwd →
+31% vs. prior week
95% CI: +24% to +38%
Causal Attribution
Weather (cold front)+22%
Baseline trend+6%
Promotions+3%
Competition0%
Ad Timing Signal
Pause ads this week.
71% of demand is weather-driven. Paid traffic would be wasted. Resume Mar 31.
The Market

We're building a shared data network
for retail demand intelligence.

Think IMS Health for retail. Every brand that joins makes the model more accurate for every other brand. WeatherVane replaces a brand's patchwork of inventory tools, ad platforms, and forecasting software with one demand intelligence output — powered by the pooled corpus across all brands.

TAM $42B SAM $3.8B SOM $100M
TAM — $42B
Demand forecasting + inventory management + retail analytics software globally (IDC 2024).
~450K retail/CPG brands x ~$93K avg annual spend
SAM — $3.8B
Mid-market DTC and omnichannel brands ($5M-$500M revenue) most underserved by enterprise tools and most likely to benefit from cross-brand intelligence.
~38K brands x $100K/yr · Most underserved by Oracle/SAP
THREE-YEAR ROADMAP
Yr 1 10–20 design partners across 3 verticals. Prove the cross-brand model. Build the corpus.
Yr 2 Commercial launch. 100+ brands on the platform. First $1M+ ARR. Open the API to ad platforms and inventory systems.
Yr 3 1,000+ brands. The standard for demand intelligence. $10M+ ARR. The corpus becomes the moat — no one can replicate 3 years of pooled data.
The Competition

What a brand gets from us
vs. what exists today.

The core difference: we pool data across brands. Everyone else analyzes one brand at a time. That means we can answer questions no single-brand tool ever could.

Capability Legacy
Oracle, SAP
Point Solutions
Planalytics, APIs
Foundation Models
TimeGPT, Chronos, TimesFM
WeatherVane
Tells you why demand moved
Learns from other brands' data no (single-series)
Separates weather, promos, competition correlation only
Adapts to your specific brand config only
Gets smarter as more brands join
Mid-market DTC pricing
Tells you when to pause or push ads index only
Confidence ranges, not point guesses partial
Works on day one (no cold start)
Your data stays private from competitors
Why nobody else can do this: We have commercial data-sharing relationships with real brands. OpenAI can build a great time-series model, but they don't have NorSari's order history or Homesick's 8 years of SKU-level data. The corpus takes years of relationships to assemble.
IMS Health did exactly this for pharma — pooled pharmacy data no chain would share with a rival, turned it into intelligence, and built a $50B company (now IQVIA). Retail has the same coordination problem. We're the neutral aggregator.
The Team
Nathaniel Schmiedehaus
Nathaniel Schmiedehaus
Founder & CEO, Thagorus PBC

Ten years as an operator watching demand move
with no tool capable of explaining why.

BUILT & EXITED
Homesick Candles
Zero to 8-figure revenue. Sold to Win Brands Group. 8+ years of SKU-level demand data.
FOUNDED & OPERATING
NorSari
Ski/outdoor apparel. Design partner since 2017. 44% CAGR over 8 years with this modeling. The proving ground for WeatherVane.
BUILT THE PROOF
+16% accuracy, 4x M5 threshold
Designed the causal inference + time-series architecture from first principles. Benchmark proof built solo.
OPERATOR ADVANTAGE
Warm intros to dozens of brands
10 years as a DTC operator. Knows the buyer, the pain, the workflow. First hires: ML engineer + enterprise sales.
The Ask

$1.25M
to lock in
the moat.

Raise
$1,250,000
Instrument
Post-money SAFE
Valuation cap
$8M post-money
Runway
18 months
Contact
nate@thagorus.com
Use of Funds
Engineering (model + data pipeline)55%
Data partnerships (brand onboarding)25%
Infrastructure (compute, storage)15%
Legal & ops5%
18-Month Milestones
NOW 1 10 partners Mo. 12 2 NeurIPS paper Mo. 15 3 Paid SaaS Mo. 18
1
Sign 10–20 design partners across 3 verticals (months 1-12)
Outdoor apparel, home fragrance, seasonal food & bev. Warm intros via NorSari and Homesick operator networks.
2
Submit 1 NeurIPS-quality paper (months 6-15)
Formal proof that cross-brand data pooling improves retail demand prediction. Published research opens the door to enterprise buyers.
3
Convert 3+ partners to paid SaaS (months 12-18)
Target: $100K ACV, 80%+ gross margin. First revenue validates the model and sets the price anchor for Series A.
GTM: Start with DTC brands in weather-sensitive verticals. Warm intros via NorSari and Homesick operator networks. The corpus compounds — earlier partners mean more data when paid accounts start.
Appendix
Honest Answers to Hard Questions
OpenAI / Google could build this
They could build a general time-series model — and several have (Chronos, TimesFM). What they can't build is the cross-brand training corpus of retail x weather x macro data, assembled through commercial relationships with real brands. OpenAI doesn't have NorSari's SKU-level transaction history. The moat is the data.
What about TimeGPT, Chronos, TimesFM?
Foundation model forecasters do zero-shot prediction on individual time series. They predict each series independently — no cross-brand learning, no causal attribution, no retail-specific training corpus. We're building the domain-specific data network they can't.
What is the +0.405 nats claim exactly?
The difference in negative log-likelihood (NLL — a standard measure of prediction quality; nats are natural units of information) between a retail-only model and one trained on diverse economic data. N=100K tokens, honest first-order result. 0.1 nats on M5 is considered significant. Ours is 4x that. Currently expanding.
Zero revenue, one design partner — too early?
NorSari has been running on the WeatherVane model since 2017 — 44% CAGR over 8 years. Homesick provides 8 years of demand data. The math works. This raise funds formalizing 10 more data-sharing relationships and growing the corpus.
Appendix
Business Model — Three Phases
Phase 1 — Now
Free Design Partners
Brands exchange data for research access. Near-zero cost to us. We want the data, not the dollars, in year one.
Target: 10 partners by month 12.
Revenue: $0 intentionally.
Phase 2 — Yr 2
SaaS Subscription ($)
Per-brand annual subscription. Weekly forecast + causal attribution dashboard + ad timing signal. Priced at a fraction of ad budget wasted per quarter.
Target: $100K ACV, 80%+ gross margin.
Initial: 3-10 paid accounts.
Phase 3 — Yr 3+
Platform API ($$)
The corpus assembled across hundreds of brands becomes training data for a domain foundation model. API access for ad platforms, inventory systems, financial data providers.
1,000+ brands. API platform open. The corpus is the product.
Year 1
$0
10–20 design partners. Build the corpus.
Year 2
$300K-$1M
3-10 paid SaaS at $100K ACV.
Year 3
$10M+
1,000+ brands. API platform. The standard.
Why Phase 1 is free: Each design partner's data is worth more to the model than the subscription revenue they'd pay in Year 1. We're buying the corpus with free access, then monetizing once the accuracy speaks for itself.