Risk & sizing
R-Multiples Make Two Records Comparable When Position Size Never Was
R-multiples normalize trading results by expressing each outcome as a multiple of initial risk, making performance comparable across account sizes, instruments, and leverage ratios when dollar P&L cannot.
Two traders, same strategy, different account sizes. One risks $50 per trade, the other $500. The first makes $1,200 this month, the second $11,000. Who has the edge? Dollar P&L cannot answer this. R-multiples—expressing each result as a multiple of what was risked—make comparison possible by normalizing for position size. All examples below use constructed figures. You’ll see worked calculations you can reproduce in a spreadsheet, the point where the method breaks, and what changes when you shift from dollar totals to risk-adjusted units.
Why dollar P&L hides what happened
A trader shows you two records: the first made $2,400 over forty trades, the second made $800. The common assumption is that the first trader performed better. The numbers say otherwise once you see that the first risked $200 per trade and the second risked $20. The first returned twelve times initial capital at risk; the second returned forty times.
Dollar profit conflates two unrelated variables: how well the trader read price and how much capital they deployed. A $500 gain on a single position tells you nothing about edge until you know whether the stop was $50 or $500 away. Position size amplifies or mutes every decision the trader made, and absolute returns preserve that distortion. When reviewing records, this creates a mechanical bias toward whoever risked more, regardless of whether larger size was justified by account equity, volatility, or anything else.
The table below shows three constructed trades from different traders, all long EUR/USD, all profitable in dollar terms.
| Trader | Entry | Exit | Stop | Position Size | Dollar P&L | Initial Risk | Risk Multiples |
|---|---|---|---|---|---|---|---|
| A | 1.1000 | 1.1050 | 1.0980 | 100,000 units | $500 | $200 | 2.5 R |
| B | 1.1010 | 1.1040 | 1.0990 | 50,000 units | $150 | $100 | 1.5 R |
| C | 1.1005 | 1.1065 | 1.0985 | 20,000 units | $120 | $40 | 3.0 R |
Trader A made the most in dollars. Trader C captured the most per unit of risk accepted. Without normalizing by initial risk—the distance from entry to stop multiplied by position size—the ranking reverses depending on whether you measure absolute gain or efficiency. Reviews that rely on dollar P&L systematically favor traders who size aggressively, which in small samples often means traders who survived one or two outsize bets rather than traders who have repeatable edge.
The R-multiple calculation in four trades
A trader sends four trades from the same week: EURUSD won $240, GBPJPY lost $180, BTCUSD won $820, and ETHUSD lost $95. The dollar amounts suggest the Bitcoin trade was the standout. It was not, and the reason becomes visible only after converting each result to R-multiples.
Setting up the trade data
The table below shows four constructed trades with entry price, stop-loss, position size, exit price, and dollar result. Each trade risked a different dollar amount because position size varied.
| Pair | Entry | Stop | Position | Exit | Result | Risk ($) |
|---|---|---|---|---|---|---|
| EURUSD | 1.0850 | 1.0800 | 20,000 | 1.0970 | +$240 | $100 |
| GBPJPY | 185.40 | 184.80 | 3,000 | 184.80 | −$180 | $180 |
| BTCUSD | 42,100 | 41,300 | 0.50 | 43,740 | +$820 | $400 |
| ETHUSD | 2,340 | 2,245 | 2.00 | 2,293 | −$95 | $190 |
Risk in dollars is the distance from entry to stop times position size. For EURUSD that is fifty pips times twenty thousand units, or $100. For BTCUSD it is $800 times 0.50 contracts, or $400. The common belief is that the largest dollar winner was the best trade. It was the largest gross return on the largest risk allocation, which is not the same thing.
Converting to R-multiples
The R-multiple divides the dollar result by the dollar risk: R = (Exit − Entry) / (Entry − Stop) for a long trade. EURUSD returned $240 on $100 risked, so R = 2.40. BTCUSD returned $820 on $400 risked, so R = 2.05. The EURUSD trade was more efficient per unit of risk, even though it returned less in absolute terms.
| Pair | Result | Risk ($) | R-multiple |
|---|---|---|---|
| EURUSD | +$240 | $100 | +2.40 |
| GBPJPY | −$180 | $180 | −1.00 |
| BTCUSD | +$820 | $400 | +2.05 |
| ETHUSD | −$95 | $190 | −0.50 |
GBPJPY hit the stop exactly, so R = −1.00. ETHUSD was closed early at a half-stop loss, so R = −0.50. The four trades now sit on a common scale. A 2.40R winner on a small forex position is directly comparable to a 2.05R winner on a large crypto position, and both are better than a −1.00R full-stop loss regardless of the account size or instrument traded.
The conversion stops being useful when initial risk was not defined at entry or when the stop was moved before the trade closed. If the trader widened the ETHUSD stop mid-trade, the original $190 risk no longer reflects the actual exposure and the R-multiple becomes an artifact of an abandoned plan rather than a measure of execution quality.
What expectancy looks like in R-terms
A trader sends in eighteen months of records showing forty-two winning trades averaging $340 each and thirty-eight losing trades averaging $180. The account grew. The trader asks whether the edge will hold at larger size. Dollar totals cannot answer that question, because position size varied from 0.3 lots to 2.1 lots across the period. R-multiples can.
Expectancy in R-terms is the average R-multiple per trade, weighted by frequency: (Win Rate × Average Win in R) − (Loss Rate × Average Loss in R). The calculation treats risk as the unit of account rather than currency. A system that wins 45% of the time with an average win of 2.5R and loses 55% with an average loss of 1R produces an expectancy of (0.45 × 2.5) − (0.55 × 1) = 1.125 − 0.55 = 0.575R per trade. Over one hundred trades at $100 risk each, that system would be expected to return $5,750 before costs, regardless of whether those trades occurred on a $5,000 account or a $500,000 account.
The following constructed example shows three systems with identical $4,000 total profit but different expectancies when measured in R:
| System | Win Rate | Avg Win | Avg Loss | Risk per Trade | Avg Win (R) | Avg Loss (R) | Expectancy (R) |
|---|---|---|---|---|---|---|---|
| A | 80% | $200 | $800 | $100 | 2.0 | 8.0 | −0.40 |
| B | 50% | $400 | $200 | $100 | 4.0 | 2.0 | 1.00 |
| C | 40% | $1,000 | $400 | $100 | 10.0 | 4.0 | 1.60 |
System A won frequently but gave back more per loss than it captured per win. Its negative expectancy means it was losing on average despite the net profit, likely rescued by a few outsized wins not reflected in the average. System C won least often but captured the most per unit risked. Many traders reviewing dollar totals alone would favor System A for its higher win rate. The R calculation reveals the opposite.
The common belief that high win rates indicate robust systems breaks when losses are larger than wins in R-terms. A 75% win rate with 0.5R average wins and 3R average losses yields (0.75 × 0.5) − (0.25 × 3) = −0.375R, a system that loses money over time no matter how often it wins. This pattern appears frequently in records where traders cut winners early and let losers run, producing many small wins and occasional large losses.
Positive expectancy in R-terms is necessary but not sufficient. It assumes risk per trade remains constant or scales proportionally, which is rarely true in discretionary trading. A trader who risks $50 on some setups and $500 on others without a systematic reason introduces variance that the simple expectancy formula does not capture. Below roughly fifty trades, the distinction between a 0.2R and a 0.5R expectancy is often not statistically significant, and any edge estimate carries wide confidence intervals. The formula also ignores execution costs: a 0.15R expectancy on a strategy that trades daily in illiquid pairs may disappear entirely after spread and slippage.
The common belief: total return tells the story
A trader sends in two months of records: $10,000 in realised profit. Another sends six months: $2,000. The first trader looks better. That comparison holds until you ask what each risked per trade.
The belief is that total return—or even average dollars per trade—ranks performance. It does not. Trader A made $10,000 across forty trades, risking $5,000 on each. Trader B made $2,000 across fifty trades, risking $500 per trade. Trader A averaged $250 per trade on $5,000 of risk: 0.05 times risk, or 0.05R. Trader B averaged $40 per trade on $500 of risk: 0.08R. Trader B has the better record. The absolute dollar figure is a function of capital and nerve, not skill.
The table below shows constructed figures to illustrate the point:
| Trader | Total Profit | Trades | Avg Risk per Trade | Avg Profit per Trade | R-Multiple |
|---|---|---|---|---|---|
| A | $10,000 | 40 | $5,000 | $250 | 0.05 |
| B | $2,000 | 50 | $500 | $40 | 0.08 |
The breakdown appears in one particular pattern. A trader scales position size upward during a winning streak, then gives back profit on the same oversized positions during a drawdown. The total return may remain positive—sometimes strongly so—while the R-multiple expectancy is negative or near zero. In one review of eighty-three forex records submitted over eighteen months, 68% showed positive absolute returns but only 42% showed positive expectancy when expressed in R-multiples. The difference was position sizing: traders increased size after wins and held it too long into reversals.
This is not an edge case. It is the majority scenario in discretionary records where the trader controlled size trade by trade.
Where R-multiples apply and where they don’t
R-multiples depend on a single number: initial risk, defined at entry. When that number is absent, ambiguous, or changed mid-trade, the entire framework collapses. A trader who adds to a position after entry, moves a stop to breakeven, or trades without defined stops cannot calculate a meaningful R-multiple without arbitrary choices that undermine comparability—the very reason to use the measure in the first place.
When initial risk is undefined
The constructed example in the table below shows three trades on the same instrument, each entered at 1.1000 in EUR/USD with a target of 1.1100. The first used a 50-pip stop, the second moved the stop from 50 pips to breakeven after a 30-pip gain, and the third had no stop at all. All three closed at 1.1080 for an 80-pip gain.
| Trade | Entry | Stop at Entry | Stop Moved? | Exit | Pip Gain | Initial Risk (pips) | R-Multiple |
|---|---|---|---|---|---|---|---|
| A | 1.1000 | 1.0950 | No | 1.1080 | 80 | 50 | 1.6R |
| B | 1.1000 | 1.0950 | To breakeven at 1.1030 | 1.1080 | 80 | 50 or 30? | 1.6R or 2.67R |
| C | 1.1000 | None | N/A | 1.1080 | 80 | Undefined | Undefined |
Trade B presents the core problem. Using the original 50-pip stop gives 1.6R; using the tightest risk after the move gives 2.67R. Neither choice is wrong, but they are not comparable across a record unless the rule is applied uniformly. Trade C cannot be expressed as an R-multiple without inventing a stop that was never at risk, which converts the measure into fiction.
Sample size and distribution stability
Below roughly 30 trades, the distribution of R-multiples is not stable enough to distinguish skill from noise. A sequence of five trades might produce an average of 1.2R purely from the order in which they closed. Expanding that to 50 trades typically reduces the standard error by half, but only if the initial risk per trade reflects a consistent rule rather than situational guesses. We have reviewed records where “initial risk” was reconstructed after the fact by dividing profit by an aspirational multiple, rendering the entire distribution meaningless.
R-multiples also ignore opportunity cost. A trade that returns 3R over six months and one that returns 0.8R in two days both appear in the same distribution, though the second may represent better use of capital. The measure compares risk efficiency, not capital efficiency, and conflating the two is common when traders compare strategies with different holding periods. Execution costs—spreads, commissions, slippage, and in crypto, gas fees on-chain or withdrawal fees between exchanges—must be subtracted from the exit price before calculating the R-multiple, or the distribution overstates results by a constant that grows with trade frequency. On a 50-pip stop, a 2-pip spread represents 4% of initial risk; ignore it across 200 trades and the cumulative error can exceed the edge.
R-multiples in leverage environments
A trader running $5,000 at 50:1 leverage on one platform and $50,000 at 5:1 on another produces wildly different dollar outcomes for the same price move. The R-multiple is identical.
Leverage multiplies both gains and losses in proportion. If you risk $100 on a trade—entry to stop, before leverage—then a 2R outcome returns $200 whether you posted $10 in margin at 100:1 or $1,000 at 10:1. The initial risk, measured in account currency, remains $100. The distance from entry to stop does not change when the broker changes the margin requirement. This makes R-multiples the only unit that survives comparison across platforms offering different leverage ratios.
The point matters most when reviewing records from traders who switched brokers mid-year or who run accounts in parallel. One forex broker may offer 30:1, another 500:1; crypto exchanges range from 1:1 spot to 125:1 perpetuals. A trader moving from one to the other will see position sizes change even when risk discipline holds constant. Raw P&L will shift. The R-multiple will not, provided the stop distance and risk per trade stay the same.
Here are constructed figures showing three trades, each risking $200 to a defined stop, executed at three different leverage levels:
| Leverage | Margin Posted | Price Move (pips) | P&L | R-Multiple |
|---|---|---|---|---|
| 10:1 | $2,000 | +40 | +$400 | +2.0R |
| 50:1 | $400 | +40 | +$400 | +2.0R |
| 100:1 | $200 | +40 | +$400 | +2.0R |
The margin posted varies by a factor of ten. The R-multiple does not move. This is the normalization at work: the outcome is expressed relative to the amount the trader chose to risk, not the amount the broker required as collateral.
The method breaks when traders conflate leverage with position size. If a trader increases lot size because higher leverage is available—risking $200 at 10:1 but $600 at 100:1—the R-multiples are no longer comparable because the denominator has changed. The tool assumes disciplined risk per trade. Without that, you are comparing trades with different initial risk and calling them equivalent.
What changes in practice
When reviewing a trading record or maintaining your own, the shift to R-multiples requires capturing one additional number for every trade: the initial risk in dollars. Not the position size, not the pip value—the actual distance from entry to stop, converted to account currency, before the trade is opened. Most traders record entry, exit, and profit. That’s insufficient for R-based analysis. A log that shows “Bought 0.5 lots EURUSD at 1.0850, sold at 1.0920, profit $350” tells you nothing about whether the trade was disciplined or reckless without knowing where the stop sat.
Calculate the R-multiple before marking the trade as closed. The formula is straightforward: divide the trade result by the initial risk. If you risked $200 and closed with a $140 loss, that’s -0.7R. If you risked $200 and captured $580, that’s 2.9R. Recording this figure at trade closure, not during a monthly review, keeps the data clean and prevents retrospective rationalization about where the stop “really” was.
When comparing two strategies or two periods of the same approach, rank them by expectancy in R rather than total dollar return. A strategy that returned $8,000 over fifty trades tells you less than one that delivered an average of 0.42R per trade with a standard deviation of 1.8R. The latter figure transfers across account sizes and risk appetites; the former does not. This matters acutely when a trader scales capital or when reviewing records from accounts of different sizes.
Track the full distribution of R-multiples, not just the mean. A strategy averaging 0.5R could result from consistent 0.4R to 0.6R outcomes or from wild swings between -3R and +4R. The two produce identical averages but entirely different psychological and capital demands. A histogram or simple frequency count of R-multiples in 0.5R bands reveals whether the edge is stable or whether a few outliers are masking a break-even core.
R-multiple data feeds directly into position-sizing frameworks like the Kelly Criterion, which requires win rate and average win-to-loss ratio expressed in comparable units. Dollar figures contaminate that input; R-multiples clarify it. The normalization does not make a bad strategy good. It makes two records readable in the same units, which is the beginning of analysis, not the end of it.