Skip to main content

Reliability and safety

Automated trading fails in boring ways. A webhook arrives twice. A server restarts in the middle of an order. An exchange times out after it has already accepted the order. A stop-loss is refused. This page lists the failure modes OptAlgo is built to handle, and the mechanism behind each one. It also says plainly where a mechanism only alerts the OptAlgo operations team and does not change your position.

How a signal moves through the platform:

  1. The signal listener receives your webhook, checks it (see Signal Flow) and queues it.
  2. The execution service takes it from the queue and places the orders on your exchange.
  3. The scheduler runs background jobs that compare what the exchange holds with what OptAlgo recorded.

How the platform runs​

The facts below are taken from the production deployment manifests and the service code. OptAlgo does not publish an uptime figure, and this page does not claim one.

ServiceWhat it doesHow it runs
API (api.optalgo.com, the app and the OptAlgo API)Accounts, bots, trades, the v1 API2 replicas
Signal listenerReceives webhooks and API signals, checks them, queues them2 replicas; readiness and liveness probes; rolling updates keep at least one pod serving (maxUnavailable: 0); 45-second shutdown grace so in-flight webhooks finish
Execution service (worker)Places and protects orders2 replicas; a liveness probe restarts a pod whose queue consumer has been dead for about 90 seconds plus up to four 30-second checks; a 150-second grace period lets a draining pod finish running signals (up to 2 minutes) and hand the rest back to the queue
Signal queueCarries signals from the listener to the workerRabbitMQ on persistent storage (50 GiB); the queue and its messages are durable; a message is acknowledged only after its signal has been executed
RedisPer-bot locks, counters, cachesAppend-only persistence (appendfsync everysec) on top of snapshots: a restart loses at most about a second of writes; eviction is off, so a lock is never dropped for memory
SchedulerThe reconciliation jobs below1 replica
Strategy engine (OptAlgo's own strategies)Computes system-strategy signals2 pods on the same feed; only the lease holder acts, the other takes over on a crash or deploy; a disruption budget keeps one pod available. It does not run My Bots

Deploys are GitOps: every service's image tag is pinned in git and applied by ArgoCD with automated sync, prune and self-heal. Rolling back is a git revert of the tag; a manual change on the cluster is reverted automatically.


Security​

  • API keys (the OptAlgo API) are shown once at creation; OptAlgo stores only their SHA-256 hash and a short display prefix. Keys carry scopes (read, bots, trade; only trade reaches real money: live signals, switching a bot to live, live bots), can be revoked in the app and stop working at once on every server. The plan is re-checked on every request, and a key of a closed account answers 401.
  • Exchange API keys are encrypted at rest with RSA-OAEP (2048-bit key, SHA-256) and decrypted only inside the execution path. No endpoint returns a stored exchange key: the app and the API show the exchange, status, leverage and margin mode of a connection and nothing else.
  • Binance is connected through Binance's own OAuth consent screen (OptAlgo is an official Binance partner). You approve on binance.com; Binance creates an API key named Optalgo, restricted to OptAlgo's outbound IP address, with trading permission for spot and futures (no withdrawal permission is requested), and hands it to OptAlgo encrypted to OptAlgo's RSA public key. You never type or paste an API key. The authorize link's state is a signed token that expires after 15 minutes and names the account that started the flow; a tampered, expired or foreign link creates nothing.
  • Every API call is logged with the key that made it: route, status, latency, the bot and signal ids it touched, client IP and user agent, and for writes the request body with every secret redacted (kept 90 days). Every change to a bot made through the API appears in the bot's own log feed as … via API key "<name>" (<prefix>).
  • Telegram notice to you on every money-moving API write that is not a trade: a changed allocation, mode or max positions, a live bot created, a bot stopped (with how many positions it closes), deleted or converted, a ticker's weight, pause or removal, and a Binance connection made. Trades themselves are reported by the execution service. Turn Telegram notifications on in the app.
  • Per-key autonomy and limits. A key either makes its agent ask before live trades (confirm_live, the default) or lets it trade on its own (full). Any key can carry limits that the server enforces on live bots: maximum allocation per bot and maximum leverage (both required for full), allowed exchanges and account types, and a daily loss limit that pauses the key's new live risk until the owner resumes it (exits, closes, stops and paper bots still work; deleting a losing bot does not reset the count). A request beyond a limit is refused before it reaches the exchange and the owner is told on Telegram; if today's PnL cannot be read, the live entry is refused rather than let through. Only the owner, logged in to the app, can change autonomy and limits; an API key gets 401 on that endpoint.
  • Webhook keys (a bot's strategy_key) trade the bot with no scope check, live included, so the API returns them only to keys with the trade scope.

Receiving and queuing signals​

Every signal leaves a trace​

What can go wrong: a signal "did nothing", and nobody can tell whether it was rejected, lost or failed on the exchange.

What OptAlgo does: each signal gets a correlation id and one row in an internal signal ledger. The listener records its decision: processed, skipped or rejected, with the reason. The execution service records every attempt on the same row: start and finish time, duration, outcome and reason, and how often it was retried. Rejections and errors are also written to the bot's logs and notifications in the app, so you can see why a signal was refused. The ledger itself is internal; the OptAlgo team uses it to answer support questions precisely.

The queue survives restarts​

What can go wrong: the execution service restarts (a deploy, a crash, a node problem) while signals are waiting or running.

What OptAlgo does:

  • Signals are published to a durable queue as persistent messages.
  • A message is acknowledged only after its task has finished. A task interrupted by a crash is delivered again instead of being lost.
  • On a planned shutdown, the execution service drains. It stops taking new messages, hands back the ones it had not started, and gives the running ones up to two minutes to finish before it exits.

Temporary database outages​

What can go wrong: the cache that holds the per-bot locks is briefly unreachable.

What OptAlgo does: the signal is held, unacknowledged, until the cache answers again. It is not executed without its lock. Entries are held only while they are still fresh. Closes are held for up to 20 minutes.

Stale signals are not executed​

What can go wrong: a signal is delayed and would act on a market that no longer reflects the alert, for example open a position at a price that has moved on.

What OptAlgo does: signals older than a freshness limit are dropped and the delay is logged. Full closes (CLOSE, FLAT) are the exception: they are never dropped for being old, because dropping an exit would leave the position open.

One signal at a time per bot​

What can go wrong: two signals for the same bot run at the same time. For example, a close starts while the entry is still placing its stop-loss.

What OptAlgo does: each bot (and each trade_id within a bot) has a lock. A signal that finds the lock taken waits and is retried about once a second, so signals for the same bot run one after another. A lock held by a long-running task is renewed while the task runs. If a signal cannot get the lock within about four minutes, it is captured as a dead letter (see below) instead of running out of order.

Duplicate deliveries​

What can go wrong: the queue delivers the same message twice after a restart, so a partial close or a stop move would run twice.

What OptAlgo does:

  • Entries: an entry is refused when the bot already has an open trade (There is an already open trade).
  • Closes: a close is refused when there is no open trade.
  • Other actions: partial closes, stop moves, replace signals (UPDATE_ORDER, CHANGE_DIRECTION), chase orders, order cancels and tasks are marked as done after they succeed. A redelivered copy of the same message is skipped for an hour.
Webhook signals are not deduplicated by transaction_id

Each webhook POST is a new signal. If TradingView or your script sends the same alert twice, the second copy is processed on its own. A second LONG is refused because the trade is already open. A second CLOSE_PARTIAL, however, closes another slice. OptAlgo does not use transaction_id to drop webhook signals, because many TradingView setups reuse one static transaction_id for every trade.

Dead letters​

What can go wrong: a signal fails in a way nobody anticipated.

What OptAlgo does: the signal is captured as a dead letter, with its error, its context and its full payload (exchange credentials are removed), and the operations team is alerted. The team can replay it once, after safety checks: an entry only while it is still fresh and has no trade yet, an exit only while the trade is still open. Replay is a manual, operator-only action.


Talking to the exchange​

Timeouts after the exchange accepted the order​

What can go wrong: the exchange accepts an order but the response is lost. A blind retry would open a second position.

What OptAlgo does: every order carries a client order id. After a recoverable error, the execution service first looks the order up by that id, and only places it again if the exchange really has no such order.

Retries for temporary failures​

What can go wrong: a network error or a temporary exchange error interrupts an order.

What OptAlgo does:

  • On Binance, Bybit and OKX, a failed execution is retried up to 5 times, with a short delay that grows on each retry.
  • A close gets one extra forced retry.
  • An entry that the exchange definitively rejected (for example for insufficient balance) is not retried.

Errors are classified​

What can go wrong: retrying an error that can never succeed. A wrong API key produces 3 to 5 failed attempts and noise instead of one clear message.

What OptAlgo does:

  • Authentication errors are not retried. You get one notification asking you to check the API key, its permissions and the IP whitelist.
  • Exchange rate limits pause entries for that exchange account for the time the exchange asks for (one minute when it does not say, at most 15 minutes). Exits always run.

Fast entries​

What can go wrong: entries are slow because each one fetches the balance, the price and the account settings one after another.

What OptAlgo does: the balance and the price are fetched in parallel. When many accounts enter the same market at once, they share one price lookup. Binance futures account settings (position mode, multi-asset mode, margin mode) are cached.

Precision and minimum size​

Quantities and prices are rounded to the exchange's precision. An entry is checked against the exchange's minimum order value before it is sent. Leftover dust after a close is treated as closed.

Several bots on one account​

What can go wrong: two bots, or a bot and a manual trade, share one exchange account. A cleanup meant for one bot cancels the other's orders, or a close for one bot touches a position the other bot (or you) opened.

What OptAlgo does:

  • Every order OptAlgo places carries a client order id with OptAlgo's broker prefix for that venue. The cleanup paths (orphan order cleanup, the cancels of a full flush) act only on orders that carry the prefix or are referenced by the trade being flushed. An order you placed by hand is never cancelled.
  • Hedge mode. Binance USDⓈ-M futures and OKX are traded in hedge mode: before an entry, OptAlgo reads the account's position mode and switches it to hedge mode if it is in one-way mode, and every order carries its position side (positionSide on Binance, positionIdx on Bybit). A long and a short on the same symbol are separate positions, and a bot's close touches only its own side. If Binance refuses the switch because a position is open in one-way mode, the entry fails and you are emailed (once a day) to close that position and switch to hedge mode.
  • Each bot has its own lock, its own trade records and its own allocation, so five or six bots on one Binance account is the normal case.

Symbols and market types​

Each exchange and account type has a curated symbol list (GET /v1/exchanges/{exchange_id}/symbols through the API, the symbol picker in the app). A futures ticker ends in .P (BTCUSDT.P), a spot ticker does not (BTCUSDT), and the suffix decides the market type everywhere: a signal whose suffix does not match the bot's market type is rejected before anything is sent.


Protecting positions​

No position without a stop-loss​

What can go wrong: the entry fills, but the stop-loss order is refused, and the position sits on the exchange without protection.

What OptAlgo does:

  • If a signal asked for a stop-loss and none could be placed after the entry filled, OptAlgo closes the position.
  • If the entry has not filled yet, the order is left resting and you are alerted that it will have no stop-loss if it fills.
  • When a stop is moved (MOVE_STOP_LOSS) and the new stop cannot be placed, the position is closed.
  • After a partial close, the stop-loss and take-profit are placed again for the remaining quantity. If that fails, you are alerted.

The trade record is always written once the entry is on the exchange, even if the stop-loss step fails. A later signal can then still close the position.

Binance conditional orders​

What can go wrong: Binance USDⓈ-M futures keeps stop-loss and take-profit orders in a separate "algo order" service. Ordinary open-order lists do not show them, and the ordinary cancel call does not reach them. A system that ignores this cannot see or cancel its own stops, or may cancel live stops by mistake.

What OptAlgo does: it places, reads and cancels conditional orders through the algo service. It stores their ids on the trade, and it counts them as the trade's own orders. Background cleanup never cancels an order that belongs to an open trade.

Exits always get through​

What can go wrong: you stop a bot, pause a subscription or your plan expires while a position is open. A check that blocks new trades would then also block the exit or the stop move.

What OptAlgo does: when an open trade exists, exits and stop moves pass those checks (the exit pass). UPDATE_ORDER and CHANGE_DIRECTION are reduced to their closing half. New entries stay blocked.

Exits follow the trade's venue​

What can go wrong: a bot is switched between paper and live while a trade is open. The exit would then go to the wrong venue: a real order for a paper position, or a "close" on paper while the real position stays open.

What OptAlgo does: an exit always runs on the venue (paper or live) its trade was opened on.

Multi-symbol position slots​

What can go wrong: two coins of a multi-symbol bot race for the last free slot, and both open.

What OptAlgo does: a slot is reserved atomically before the entry is queued. Only one of the racing coins gets it; the other entry is rejected with a notification. A reservation that does not lead to a trade is released, or expires after 20 minutes. The execution service checks the bot's max positions again, under the bot's lock, before it opens.


Background reconciliation​

The scheduler runs these jobs continuously. "Acts" means the job changes orders or positions; "alert-only" means it notifies the OptAlgo operations team and changes nothing.

JobEveryWhat it checksWhat it does
Orphan order cleanup5 minOpen orders placed by OptAlgo that belong to no open trade. Binance conditional orders are included.Acts. Re-checks after 60 seconds, then cancels the orphan. Orders of open trades are never touched.
Safety-check heartbeat5 minBots with a safety check whose VALIDATE_ALERT heartbeat has been missing for twice the configured interval while a trade is open.Acts. Sends a CLOSE for the open trade.
Copy-trading consistency5 minFor subscriptions to OptAlgo strategies: a subscriber's trade still open while the strategy's own trade has closed.Acts. Re-checks after 60 seconds, then closes the subscriber's trade.
Order state sync15 minFor subscriptions to OptAlgo strategies on Binance futures: compares the quantities of filled orders on the exchange with the trade records (1% tolerance).Acts. Refreshes the trade from the exchange, which records stop-loss or take-profit fills.
Quantity reconciliationabout every 15 minFor subscriptions to OptAlgo strategies: compares exchange positions and resting orders with the trade records, per symbol and side (0.6% tolerance).Alert-only. Confirmed by a second run after 60 seconds. Any fix is applied by an operator.
Position snapshots6 hRecords each connected account's positions.Baseline for quantity reconciliation.
Unprotected position watchdog5 minLive trades with a stop-loss price that have no live stop order on the exchange (Binance).Alert-only.
Untracked position watchdog5 minFutures positions worth more than 10 USD that match no open OptAlgo trade.Alert-only.
Lost signal watchdog5 minSignals the listener queued that no execution attempt picked up, or that failed without a retry, plus new dead letters.Alert-only. Never replays anything.
Queue and service health3 minSignal queue without consumers or with a growing backlog, and the health of the listener and signal stream.Alert-only.
Weekly PnL reportSunday 22:00 UTCYour realized PnL for the day, the last 7 and 30 days, and in total.Sent by Telegram to users who enabled Telegram notifications.

Some checks currently cover only part of the platform: the copy-trading, order-sync and quantity jobs look at subscriptions to OptAlgo strategies and skip some exchanges, and the watchdogs check a limited number of accounts per run, in rotation. They add a safety net. They do not replace the protections in the execution path described above.


Known limits​

  • A 204 reply only means "received". The webhook endpoint answers 204 even when a check rejects the signal, and it cannot ask the sender to retry. Check the bot's logs to confirm what happened. Signals sent through the OptAlgo API get a transaction_id, and GET /v1/signals/{transaction_id} returns the outcome, including refused when the executor declined a duplicate entry or a close with nothing open.
  • Paper bots have no exchange-side stop-loss or take-profit. Paper trades open and close on your signals. Send a CLOSE when your stop or target is hit.
  • Margin accounts: new entries on Binance margin accounts are currently disabled. Exits still run.
  • Watchdogs observe; they do not trade. The alert-only jobs above tell the operations team about a problem. Fixing it is a human decision.