How we rank futures signals
Five tests, each applied the same way to every service. A test is marked passed only when a buyer could reproduce it without taking the provider's word for anything.
The ranking logic is deliberately blunt: tally how many of the five tests a service clears outright, and where two finish level, let the weight of the partial evidence break the tie. No commission tilts the ranking, and no amount of money can buy a place in it. The whole point is to reward what can be checked over what is merely posted — which is why a service with an unglamorous record you are free to inspect finishes ahead of a dazzling one you would simply have to trust.
The five tests
1. Locked before the outcome was known
Each call is hashed and written to a public ledger at the moment it is published, so its entry, target, stop and grade cannot be edited, re-priced or back-dated once the contract resolves.
2. A track record you can re-run
A continuous, real-money history corroborated by an outside party, shown with return, drawdown and win rate — not a curated reel of green trades with the red ones quietly edited out.
3. Conviction grades with arithmetic behind them
An A-to-D label on every call, set by where it sits in that model's own return distribution, instead of a mood word like “strong setup” that means whatever the sender wants.
4. Pricing on a public page
Every cost and every trial term visible before a buyer is asked for an email or a card — no “message us for rates”.
5. Revenue that isn't the click
Income that comes from the subscription itself, not from broker affiliate kickbacks that quietly reward the volume of sign-ups over the quality of the signal.
The same five tests, run against the field
Applied identically, the tests sort the market into types rather than individual brands. The grid below is the rubric laid over the archetypes a futures buyer actually meets — the chat channel, the copy-trading room, the social caller, the aggregator — set beside the corroborated desk. The point is not that the top pick is praised more loudly; it is that its column is the only one with every box ticked.
Read it down a column, not across a row. The test almost nothing clears is fixed before settlement, which is why it leads the rubric: a service can hold a genuinely good record and still fail it, simply because the record was never committed anywhere a stranger can re-check.
A win rate is nothing without its denominator
On its own a percentage is a headline rather than evidence. Print “95% win” with no count beside it and the figure could be nineteen of twenty hand-picked tickets, or it could be hiding every losing week behind the wins; the number itself gives you no way to separate the two readings, and that ambiguity is the whole reason a seller leans on it.
Compare how the desk pick states the same kind of figure: 74.4% across 78 Swing Trade signals in 2026. That 78 is the denominator — the complete tally of calls, losers and all, over an unbroken stretch. Supply it and the percentage becomes interrogable: about 58 of the 78 settled green while the remainder did not, and the +225% then reads against a real drawdown instead of drifting free of one. Given a choice, a smaller win rate that arrives with its count tends to be worth more than a larger one that arrives without, since the tally is the single figure a dishonest desk cannot inflate short of an outright lie.
Before any win rate earns your trust, press it on two points — over what total, and are the losing calls inside that total? A figure that cannot answer both belongs in the marketing column, not the record.
What a conviction grade has to mean
What the third test demands is a grade that is computed rather than picked. On the desk pick that grade is fixed per model, drawn from each model's own measured returns, which lets it hold up when read across very different horizons:
| Model | Horizon | Grade-A bar (per trade) |
|---|---|---|
| Day Trade | intraday, opened and closed inside one session | 0.70% avg / trade |
| Multi Hour | from a few hours out to a couple of sessions | 4.50% avg / trade |
| Swing Trade | roughly one to four weeks per position | 6.00% avg / trade |
| Investing | long-horizon, highest-conviction calls | long-horizon |
An A marks the top band of a model's own measured return spread; a D is the lowest band still published. Because the bar is set per model, an A on an intraday call (a move near 0.70%) and an A on a swing call (nearer 6.00%) both read as “top band for this horizon” rather than one absolute figure forced across very different holding times. There is no E grade — it was retired from the live product so the four-step ladder keeps its meaning.
The per-model bar is also why the four-model book matters even if you only follow one horizon: each grade is calibrated against its own model's spread, not flattened against a slower model's far larger swings. One yardstick stretched across every model would make every fast call read as weak and every long-horizon call read as strong, which would tell a buyer nothing.
Why the timestamp test sits at the top
A win-rate banner is the cheapest thing in the world to manufacture; a public, pre-outcome timestamp on every single call is one of the hardest. The rare combination that closes the door on retroactive editing is an independently corroborated multi-year record and a per-call cryptographic receipt. As of 2026 the only service in this guide passing all five tests is the #1-ranked provider. We grade it on the structure of its record alone and make no assertion about which contracts or markets its models trade. How the timestamp works, and how you check one yourself, is set out on the timestamping test and the verification walkthrough.