LongShortNMargin Research software · Founded 2025 · Vietnam Status · Pre-release
Strategy test bench · Built with Claude Code

A test bench for AI agents that build trading strategies.

Agent claims, checked against the platform's own report.

LongShortNMargin is an early, founder-led software company. Its product is a test bench: Claude Code agents write a strategy spec, build the MetaTrader 5 expert advisor and run it in the platform's own tester. The result is checked against prop-firm evaluation rules (simulated accounts with a profit target and loss limits) and against an acceptance bar fixed before the run. By design, a separate session writes the verdict. No Claude agent places orders. The bench is in a private repository. The public code is mart-forge, the open-source side of the company's data tooling.

Record · 10 Oct 2026

Company record

Company
LongShortNMargin. Spoken form: Long Short & Margin.
What
Research and evaluation software. A test bench on which Claude agents research, build and test automated MetaTrader 5 strategies against prop-firm evaluation rules.
Built with
Claude Code, the primary builder. Architect, worker, analyst and reviewer sessions.
Founder
Long Vu Duc, Founder. linkedin.com/in/vuduclong0309 Public
Founded
2025
Location
Vietnam
Stage
Bootstrapped. Founder-led, one-person company. Pre-release. No customers yet.
Industry
Software. Developer tools and research infrastructure for systematic trading.

Proof

Public

Open source

mart-forge, a Claude Code plugin pack

4 plugins, 21 skills, MIT licence. 300 offline tests pass (linters and static structure, not the skills end to end). Pre-release.

github.com/LongShortNMargin/mart-forge

Private

Claude use

Claude Code is the primary builder

Architect, worker, analyst and reviewer sessions. Five Claude models recorded in the worker logs, July and August 2026.

Roles and models

Private

Test bench

Built June to August 2026. Checks re-run on 10 October 2026.

About 4,800 lines of Python and 660 lines of PowerShell, in a private repository. About 160 unique real-tick tester reports across about 25 in-house strategy families on file. These are counts of artifacts, not a success rate.

How it works · The four checks

Public

Founder's result

Passed the FTMO Challenge in six weeks.

Certificate dated 6 August 2025.

View the certificate

First phase of FTMO's two-step evaluation, on a simulated account. The certificate names Long Vu, the founder (full name Long Vu Duc). Six weeks is the founder's figure. The certificate carries the date and no duration. The pass predates the company's use of Claude. Past performance, actual or simulated, does not indicate future results.

How to read this site

Three tags. Most of the work is private. The site marks which is which.

  • Public You can open it now. The link sits beside the claim.
  • Private On file in a private repository. Described in plain words. There is no outside link to check.
  • Planned Not built yet.
Built with Claude CodePrivate

From spec to verdict.

Claude Code is the primary builder, and all Claude use so far has run through it. The work moves through six steps. By design, the session that runs a test does not write its verdict. Claude Code is not the only builder: other AI coding agents were also used on parts of the work.

  1. 01 · Spec

    Architect agent

    Writes

    The spec, before the run.

  2. 02 · Build

    Worker agents

    Write

    The expert advisor.

  3. 03 · Run

    Worker agents, through the harness

    Run

    The MetaTrader 5 tester, on real ticks.

    By design, writes no verdict

  4. 04 · Read

    Harness (code, not an agent)

    Writes

    The result manifest, after reading MetaTrader's own tester report.

    Computes no verdict

  5. 05 · Judge

    Analyst agent

    Writes

    The verdict, against the bar in the spec.

    By design, a different session

  6. 06 · Decide

    Founder

    Decides

    Anything that touches an account.

Account boundary No Claude agent places orders or attaches an expert advisor to an account. The founder attaches any expert advisor himself.

Reviewer agents at steps 01 and 05

Fresh context.

They attack the spec and the result and recompute the result from the raw report.

Verifier subagent at step 02

Read-only.

It audits diffs.

This is the design, not an absolute. The record holds catches and lapses. The method page names the catches and states that lapses are recorded.
Models
Recorded in the worker logs, July and August 2026: claude-sonnet-5, claude-opus-5, claude-sonnet-4-6, claude-opus-4-8, claude-haiku-4-5
API
Hooks and headless runs (claude -p) tie the sessions together. No in-house code calls the Claude API yet. Four API workloads are planned. They are listed on the method page. Roles, models and planned work
The test benchPrivate

Five parts. One question.

The question about any automated MetaTrader 5 strategy: would it have stayed inside a prop firm's evaluation rules, and is that better than luck? Built June to August 2026. About 4,800 lines of Python and 660 lines of PowerShell, in a private repository.

01
Tester runner. Runs the MetaTrader 5 Strategy Tester headless. A run is marked fit to be judged only when the tester log proves real ticks and history quality of at least 95 percent.
02
Report parser. Reads MetaTrader's own tester report.
03
Spec loop. A YAML spec states the hypothesis, the exact inputs changed, the declared number of trials and the acceptance bar before the run. The loop writes a result manifest and computes no verdict.
04
Rule simulator. Replays a trade sequence against a 5 percent daily and 10 percent total loss floor on floating equity, then compares the outcome with a zero-edge luck baseline.
05
Trade cards. Renders one card per trade for reviewing individual trades.

Read the method

Status · Pre-release

Early. Stated plainly.

What exists today, sorted by whether an outsider can check it.

Public

mart-forge on GitHub: MIT licence, 4 plugins, 21 skills. The founder's FTMO Challenge certificate, dated 6 August 2025, with its link and details in the proof row above. The founder's LinkedIn profile and two publications.

Private

The test bench. About 160 unique real-tick tester reports across about 25 in-house strategy families, counted as artifacts, never as a success rate. A market-data warehouse of gold (XAUUSD) price bars: about 1.24 million bar rows, 2 January 2024 to 9 July 2026, zero duplicate keys. It is not refreshed on a schedule and the data is not redistributed.

Planned

A small public extract of the bench (report parser, rule simulator with self-test, spec loop with tests) running on synthetic bars. Four Claude API workloads: reviewer agents on every spec and verdict, executor sessions on the Windows test rig, an agent that re-reads the prop firm's published rules and flags drift against the encoded rule table, and the 21 mart-forge skill specs run as an eval suite across Claude models.

Not claimed: customers, revenue, a team, or any trading result beyond the certificate.

Boundaries

What the company does not do.

No orders by Claude agents.

No Claude agent places orders or attaches an expert advisor to an account. The founder attaches any expert advisor himself.

No money managed.

LongShortNMargin is not a broker, investment adviser or asset manager. It does not manage money, sell signals or trade for anyone else.

No advice.

Nothing on this site is investment advice or an offer to buy or sell any financial instrument.

No trading results claimed beyond the certificate.

No trading result is claimed beyond the founder's certificate. Evaluation and trial accounts are simulated. Trading leveraged products carries a high risk of loss.

Founder and contact

Long Vu Duc, Founder.

Data engineer since 2019. M.Sc. Financial Engineering, WorldQuant University. B.Sc. Computer Science, Nanyang Technological University. Two publications. He reviews what the agents produce and decides anything that touches an account.