Skip to content
python.financial

Jev may help analysts classify financial text and use the results as inputs to a forecasting model. The public evidence reviewed here does not show that it can consistently produce profitable trades. Fast responses and probability estimates can be useful, but decisions based on them may still lose money after trading costs.

TypeSafe AI introduced Jev on September 15, 2026. The model is designed to answer defined questions across different subjects. It was not developed specifically to forecast financial markets. Developers have already connected it to trading applications, but each application also has its own trading rules and execution code. Its results depend on all of these parts. Official announcement.

This review covers public evidence available on September 21, 2026. The trading results below are reported by the authors of the linked projects.

What Jev actually does

An application sends Jev information, called state in its API, along with questions whose possible answers are defined in advance. The information could be an announcement, a customer message, or a set of current market observations. Jev returns values and probabilities that software can read directly, without a written explanation. Each question in a request is answered separately using the same information. The application then combines those answers according to its own rules. TypeSafe introduction.

The interface has three question types:

Type What comes back Financial example
Choice One of the supplied options, a probability for each option, and a confidence value Classify an announcement as guidance raised, guidance withdrawn, or another event.
Score A rating on a scale with written descriptions of each level, probabilities for those levels, and confidence Assess how strongly a filing describes liquidity concerns.
Noul The estimated probability that the answer to a yes/no question is yes Does the announcement explicitly suspend guidance?

The request formats and returned values are documented in the Choice, Score, and Noul references.

For example, a company statement saying "we have withdrawn our full-year outlook" could be submitted with a question about changes to guidance. The chosen answer could then label that announcement in a dataset. Knowing the event category does not tell us whether the share price will fall. Investors might already expect the announcement, other news might matter more, or the market might view the withdrawal favorably.

There are several separate tasks here: understanding the statement, forecasting a return, choosing a position, and executing an order. Jev can answer a question at one of these stages. The trading program must still specify how that answer affects a position and when an order should be placed.

Why traders are interested

Analysts often need to turn written descriptions into categories or numerical inputs. A model that can do this at low cost could make it easier to compare many possible inputs without training a separate classification model for each set of categories.

At the review date, TypeSafe lists jev-1.13.0 as the current model. It accepts text, charges $0.042 per million input tokens, and does not charge for output. Tokens are the units into which the service divides the input. A request can contain up to 64,000 tokens in total, with a separate 32,000-token limit for the state plus the longest question. The listed limit is 1,200 requests per minute and may change. The name jev-latest can refer to a newer version later, so an experiment should record the version returned by the service. Model reference.

At that price, a million requests containing 1,000 billable input tokens each would cost $42 in model calls. This calculation excludes hosting, market data, repeated requests after failures, and trading costs.

TypeSafe reports that Jev was 193.6 times faster and 444.6 times cheaper in its tests of automated business tasks. These tests cover work such as invoice processing and customer service, using answers from other models as the reference. They do not test financial forecasts or compare Jev with a statistical trading model developed for a particular market. Company claims, test methods.

The company calls its training method Reinforcement Learning for Calibrated Decisions, or RLCD. It aims to make predicted probabilities match how often outcomes occur. For example, events assigned a 70% probability should happen about 70% of the time across many predictions. The official material reviewed does not provide enough detail to reproduce the training method. This remains the company's description of its training goal. Users need to check the probabilities on their own tasks. TypeSafe's explanation.

What the early trading experiments show

The projects below show several ways developers have tried using Jev. Some report predictions or simulated trades. Others mainly demonstrate software that connects a model to a trading system. They test different things, which prevents a direct ranking of their results.

Project What was reported or inspected What it establishes
Hong Kong stock forecasts 54 correct next-trading-day classifications out of 120 historical cases A small direction test reported by its author. Profitability after costs was not demonstrated.
Monad/Kuru trader Code that connects Jev to order placement. Default settings use a substitute model and simulated trades Published code that can be inspected. Results from the default demo do not measure Jev.
Hyperliquid adaptation Separate positions across cryptocurrencies, with simulated trading and authenticated order submission Order placement is described by its author. We did not verify a history of returns.
Obside futures experiment A reported 3.15% simulated-account loss over roughly one day A loss under one set of trading rules during a short observation period.
Python Jev experiment engine Artificial market data, simulated delays, and a default substitute model based on hand-written rules Software for running experiments. Results from the substitute model cannot establish how well Jev predicts prices.

A 45% hit rate needs a baseline

The Hong Kong experiment reports 30 cases each for Xiaomi, MiniMax, Pop Mart, and MIXUE. Correct classifications were 12, 17, 9, and 16 respectively: 45% overall. Daily returns between -0.3% and +0.3% are labeled flat, with larger moves labeled up or down. The author excludes financial fundamentals from historical inputs when reliable publication times are unavailable. The report also acknowledges the risk of using a current model to make predictions about historical dates. Stock experiment.

The 45% result needs a suitable comparison. Having three labels does not mean that 33.3% is the right benchmark. If one label occurs much more often, always choosing it might beat 45%. To assess the result, we need to know how often each label occurred and how simple alternatives performed on the same cases. The stocks also share dates and market conditions, so their results may be related. Four stocks do not necessarily provide four independent tests.

A trading prompt can disagree with execution

In the Kuru code we inspected, Jev chooses between buy and sell using order-book information, returns, and recent trades. Its instructions describe an immediate-or-cancel market order that crosses the spread: it trades at an available opposing quote, with any unfilled quantity canceled. Although the application recognizes a hold action, this particular question does not let the model choose it. Model source.

The order-placement code describes posting and replacing post-only limit orders, which are intended to rest on the order book rather than trade immediately. It also skips incoming blockchain blocks while it is still processing an earlier one. Execution source.

The difference between the instructions and the execution code matters. A market order and an order waiting on the book have different costs and may fill under different conditions. A direction forecast can look promising when it is made, yet the orders that eventually fill may perform poorly. The repository can change, so this observation applies to the files inspected on the review date.

Frequent trading can cost more than the model calls

The Obside author reports Jev decisions every 30 seconds across Nasdaq, Bitcoin, and gold micro futures. A simulated $100,000 account reportedly made 731 trades, lost approximately $3,150, and incurred about $1,650 in fees over roughly a day. Only 21% of trades were profitable after fees. The author supplied trading-cost information to the model. Experiment report.

These figures come from the author's simulation report, rather than an audited brokerage account. They show that a model can receive information about fees and still choose trades that lose money after those fees. The short experiment does not tell us how other Jev strategies would perform.

Other projects describe more limited aims. A TSLA moving-average example adds Jev decisions to a simple trading rule to demonstrate how the interface works. Its author does not claim to have validated a profitable strategy. A paper-only signals lab combines answers to multiple questions using threshold rules, but its market data is illustrative. Both are useful software examples. Neither establishes profitable trading performance.

Why probabilities can be misleading

Confidence is not a trade's win probability

The confidence value returned with Choice and Score summarizes how concentrated their probabilities are. It is higher when probability is concentrated on an answer and lower when it is spread across answers. Because confidence is calculated from these same probabilities, it provides no separate assurance that the answer is correct. The model can assign most of the probability to the wrong answer. Confidence documentation.

A probability also depends on the question asked. "Should this account buy?" involves expected prices, costs, preferences, existing positions, and position limits. "Will the next five-minute return exceed a defined threshold?" describes an outcome that can later be checked. The probability attached to a recommended action should not be treated as the probability of that price outcome.

Calibration varies by task

The independent ASSAY-001 report measured how closely Jev's probabilities matched observed results. Its expected calibration error was 0.0204 on CLINC150 and 0.0936 on BANKING77, two public language datasets. Lower values indicate closer agreement under this measure. The experiment used Choice questions and set an acceptance threshold of 0.05 in advance. Across 8,576 responses, it found no output-format violations after amending its tolerance for rounded probabilities. These findings apply to the datasets and setup tested. Full report and amendments.

A separate BANKING77 experiment supplied category definitions and selected relevant labeled examples for each request. It reported 92.40% classification accuracy. Its authors acknowledge that they do not know whether the model had encountered this public dataset during training. The result shows how the information supplied with a question can matter. It does not measure financial forecasting or establish that probabilities are accurate for market outcomes. Experiment using selected examples.

For trading, check whether events assigned similar probabilities happen about as often as predicted in the markets you use. Examine results by asset, forecast horizon, and market conditions. Pay particular attention to predictions that lead to actual positions. An average across all predictions may look acceptable even if probabilities are unreliable for the trades you choose to make.

Directional accuracy is not profitability

Suppose a correct prediction earns 10 basis points, an incorrect one loses 10, and the round-trip cost is 3. A 60% hit rate gives an expected result of:

0.60 × 10 − 0.40 × 10 − 3 = −1 basis point per trade

These numbers illustrate the calculation. They are not Jev measurements. The size of gains and losses, together with trading costs, determines what a given hit rate is worth. With different payoffs, a lower hit rate could be profitable. See transaction costs and slippage for how these assumptions enter a backtest.

Old data does not rule out knowledge of later events

A current model may have learned information that appeared after the date of a historical prediction. Supplying only old price bars or documents does not remove what the model already learned during training. If its training cutoff is unknown, it is difficult to rule out knowledge of later events.

Running the system on historical data can still help find software errors and compare trading rules. Stronger evidence comes from recording predictions before the events being predicted occur. This is another possible source of look-ahead bias. When checking what information was available at the prediction date, consider the model's training as well as the supplied data.

Where Jev could be useful

It makes sense to start with tasks that involve interpreting text. Their results can be checked against the source material before asking whether they help predict returns:

  • Event classification: distinguish guidance changes, financing announcements, management changes, and operational disruptions.
  • Filing analysis: classify business descriptions and identify passages relevant to an analytical question.
  • Model inputs: turn judgments about text into data columns for a conventional forecasting model.
  • Reviewing records and news: identify ambiguous records or potentially relevant news for further analysis.

TypeSafe has published examples of this kind of work. Its filing-classification example reports 39 correct industry-group labels out of 60 selected annual reports. When the model is uncertain, reporting a broader category raises the number counted as correct to 48 out of 60. That second result allows answers at different levels of detail, so it measures something different from exact industry-group accuracy. The company selected the sample, and the code identifies Jev 1.12. This is a company demonstration using an earlier version, not a market forecast. SEC classification example.

Another worked example uses Jev's answers about wine reviews as inputs to a CatBoost model that predicts review scores. It demonstrates how judgments about text can be combined with a conventional forecasting model. Whether the same approach helps predict financial returns remains a separate question. Example of generating model inputs.

In a financial study, the question would be whether these text-based inputs add useful information beyond prices, volume, and the other data already available. Compare the same forecasting model with and without the Jev inputs. A result showing no reliable improvement is still useful: it tells you that the additional model calls have not justified their cost in that test.

Use ordinary code for arithmetic. TypeSafe documents weaknesses in numerical precision, counting, comparing dates, and deriving exact quantities from Score levels. It also warns that irrelevant information and text written to mislead the model can affect answers. Calculate indicators, exposure, cash, and position sizes in code, and check timestamps there to decide whether a record was available at the required time. Use Jev for the parts that require interpreting text. Documented limitations.

What convincing evidence would require

A meaningful test should make it possible to answer these questions:

  1. What was predicted? Define the event, the assets included, how far ahead the prediction looks, the price source, and how missing data will be handled before evaluating results.
  2. What was known at the time? Save publication and receipt timestamps, the information sent to Jev, the model version, the questions, and the original responses. Use point-in-time data.
  3. What is the comparison? Include simple forecasts based on how often events occurred in earlier data, along with a suitable numerical or text model tested on the same observations. If Jev modifies an existing strategy, compare it with that strategy running unchanged.
  4. Were choices made before the test? Choose the questions, model inputs, trading thresholds, and methods for adjusting probabilities using earlier data. Do not use the final test outcomes to make those choices. See overfitting in backtesting.
  5. Could the decision have been executed? Account for response delays, answers that arrive too late to use, fees, spread, slippage, rejected orders, and realistic assumptions about which orders fill. Keep a record of failed model calls as well as successful ones.
  6. Does the result persist? Report how well probabilities match outcomes, how much trading occurs, returns after costs, exposure, and drawdowns across periods and assets. Many closely related decisions made in one day still represent a narrow range of market conditions.

Publish enough detail for someone else to reconstruct the result: settings, definitions of the comparison models, model and question versions, timestamped predictions, and the method used to calculate profit and loss. Show what happens when trading costs increase or thresholds change slightly. Finding one profitable combination of settings is not sufficient. Walk-forward evaluation can help compare results across time, as long as each decision uses only information available at that point.

Jev's low API price and defined answer formats make these experiments practical to attempt. The evidence currently gives practitioners more reason to test its ability to interpret financial text than to rely on it for direct trading decisions. Whether either use produces a profitable strategy still needs to be measured.

Further reading

Start with the official launch, model reference, known limitations, and independent calibration report.

To report an error or suggest a source, see our corrections policy.

Choose which optional services may run. You can change these settings at any time.