Start with a task, not a leaderboard
A research assistant, a signal service and an execution bot solve different problems. Compare products within a defined job before comparing brands. A tool without trading permission should not lose points for lacking order recovery; a bot should not receive credit for fluent market commentary if it cannot reconcile fills.
Use the automation decision tree to choose the category. Record the asset class, venue, jurisdiction, account tier and intended operating mode. Availability in one country or subscription plan does not establish availability elsewhere.
Freeze the conditions before running the test
- Record product version, access date, plan, operating system and configuration.
- Keep the input dataset, task instructions and expected outputs identical where the products support the same task.
- Use offline fixtures, a sandbox or paper mode first. Do not fund accounts merely to complete a review.
- Record free access, compensation and any material relationship alongside the result.
- Separate documented capabilities from observed behavior. Preserve the source and effective date for pricing and restrictions.
Six tasks with inspectable evidence
| Task | Procedure | Evidence to retain | Failure or limit |
|---|---|---|---|
| Research accuracy | Ask the same source-bound question with an answer available in the supplied documents; include one question the documents cannot answer. | Inputs, answer, exact citations, factual errors and abstentions. | Invented sources or unsupported certainty; fluency is not correctness. |
| Setup and permissions | Follow documented setup in a test environment; inspect each requested permission. | Elapsed setup time, errors, permission screenshots and whether credentials can be revoked. | Unnecessary withdrawal authority, hidden requirements or inability to remove access. |
| Reproducible backtest | Pin a simple strategy and dataset, run twice, then change only the fee assumption. | Version, data checksum, commands, trade ledger and before/after cost results. | Unexplained result changes, missing costs or inconsistent timestamps. |
| Risk enforcement | In sandbox mode, attempt an order outside the configured asset or size limit. | Configured limit, rejected request and resulting account state. | A warning without enforcement is not a hard limit. |
| Recovery | Use a controlled test interruption or documented offline fixture; reconnect and reconcile state. | Timeline, order identifiers, logs and final open-order inventory. | Duplicate orders, forgotten positions or an unexplained mismatch. |
| Export and exit | Export history and settings, stop the test process and revoke test credentials. | Readable files, shutdown behavior and revocation confirmation. | Unexportable results or access that remains active after revocation. |
Only run failure tests where you control the environment and the test is permitted. Do not disrupt an exchange, a shared service or another user's account. A failed setup is still a reportable result; it is not permission to bypass protections.
Use evidence states instead of unexplained stars
For each task, record passed, failed, partially demonstrated, not tested or not applicable. Attach the artifact that supports the state. “Not tested” is not zero quality, and “not applicable” must explain why the task falls outside that product's job.
A pass applies only to the tested conditions. One rejected oversized order does not prove every risk control works. A successful restart does not establish resilience under every outage. Report the attempted cases and the limits of the conclusion.
Compare costs on the same basis
Separate subscription, infrastructure, data, trading fees, financing and execution effects. State the billing currency, billing interval, annual-payment assumptions, usage limits and whether tax is included. A free tier with one active rule is not equivalent to an unrestricted open-source installation that requires paid hosting and maintenance.
Use the net-return worksheet to make cost assumptions visible. Do not subtract spread or slippage twice when actual fill prices already contain them. Compare the same workload, not each vendor's most flattering example.
When a recommendation is justified
A recommendation needs a distinct user job, comparable completed tests, current access and pricing checks, and an explanation of where the recommended product loses. Do not calculate one overall score across incompatible product categories.
If a numerical score is useful, publish the rubric and weights before testing, retain task-level results, and show how changing the weights changes the winner. Safety-critical failures should remain visible rather than being averaged away by convenience features. Until comparable evidence exists, publish a documented capability analysis—not a “best” ranking.
Performance claims need a separate record
Setup speed and feature coverage do not prove profitable trading. Label results as synthetic, backtested, paper, testnet or live. Include the evaluation period, starting capital, fees, benchmark, drawdown, sample size, data source and reinvestment assumptions. Explain missing fills, selection effects and survivorship bias.
The backtest-versus-live guide explains why those evidence states are not interchangeable. A cryptographic proof is likewise limited to its specified statement; see what trade proofs can establish.
Make the result auditable
Each published evaluation should identify the tested product and version, who or what executed the test, the date, steps, expected and actual results, relevant artifacts, limitations and correction contact. Redact secrets, account identifiers and unrelated personal data from logs and screenshots. Never publish API keys or recovery phrases.
Recheck material claims when versions, plans, region restrictions or controls change. Keep substantive corrections visible through the correction process; do not refresh dates merely to appear current. Our methodology separates sources, interpretation and completed tests.