qinbafrank
qinbafrank|Sep 10, 2026 11:50
After GPT-6Astra came out, RSI was brought to the table again. Astra scored 0.5126 on RSI Exam, which is 18% higher than 5.6Sol and more than half higher than 5.5 Sol. The circle immediately began to debate whether this could be considered as the model becoming stronger on its own. Astra's strength is real, with long link tasks, system improvements, and experimental runs. The report card is displayed there, but for self-improvement in the laboratory, the scoring rules are mostly determined by people. Models are very good at writing and reviewing, and it is not difficult to produce a complete summary of the structure. The difficulty is to give it feedback that cannot be fooled by language. The transaction can provide this kind of feedback perfectly. Predictions will be settled, orders will be executed, and handling fees and slippage will not disappear just because the review is well written. Even a complete reflection cannot change the gains and losses that have already occurred. So when NeoSoul places self evolution on NeoTrade, the entry point is actually very direct: First, let the agent run the chain of "decision result attribution modification retest retention" in the real market. Why is trading suitable for testing self-improvement? In the past two years, most AI trading products have been focused on assisting judgment: reading news, drawing graphics, giving advice, and ultimately making decisions by humans. Agent trading takes this matter one step further: continuously monitoring the market, forming judgments according to established strategies, and placing orders within the boundaries of funds. The problem arises. Does an agent know that it has been using an expired method for three consecutive months? After knowing, can we make changes? After making the changes, can we prove that the new method is better? NeoTrade's approach NeoTrade records are made at the time of decision-making, and a single run will string together the information seen at that time, reference content, Skill version, rule checking, orders, and transaction results. Run ID fixes this decision and allows it to be retrieved from the same context later. Memory only retains the filtered experience, while Skill keeps the pre - and post modified versions for easy comparison. Outdated information, biased probabilities, and ineffective strategies ultimately result in losses, but the reasons are completely different. If the attribution is wrong, the correct part will also be corrected together. So the record must be made at the time of the decision: what was seen, which skill was used, which rules were passed, and what order was placed in the end. Going back to make up for it afterwards can easily turn the consequences into causes. Learning permissions and funding permissions are separated: methods can be changed, risks are blocked by external rules. The changes mainly focus on harness: finding information, using memory, adjusting tools, and scheduling workflows; The upper limit of the weight remains clear and will not involve the model level. Transactions are suitable for verification because they provide quick feedback and expose the entire chain. But the market noise is loud, and a single profit or loss cannot explain the ability. There is a 70% probability that a portion of the judgment should have been missed. To check calibration, it is necessary to consider sufficient samples and longer cycles. And the NeoTrade system already has about 4 million agents and over 11 million predictions, which is already a scale capable of statistical verification. From a personal perspective, what NeoTrade wants to do is to replace the rater with a market (predictive market, secondary market), whether the agent can remain stable and strong, and leave evidence behind. The version that wins in Shadow Run, when converted to real money, still has the advantage, and this cycle can only stand.
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads