KrFinBERT's 0.242% daily gross spread leaves about 0.012% after Korea's explicit trading costs. The figure measures signal quality more than tradable profit, as Yang acknowledges in the section reporting it.
Round-trip explicit cost in Korea over 2023 and 2024 was roughly 0.23%: a 0.20% securities transaction tax on sells plus 0.03% brokerage. Almost nothing remains after that subtraction. Yang writes that the raw returns "are difficult to realize in practice after accounting for realistic trading frictions." With bid-ask spread and market impact added, he says, "the net profitability of the strategy may be significantly diminished and could potentially disappear." His conclusion calls the findings "consistent with a relatively efficient Korean stock market."
The abstract leaves out this arithmetic. It closes by describing LLM-based textual analysis as a powerful tool for financial decisions. A reader who stops there encounters a different paper from the one described in the portfolio section.
The English-text substitute we can run
We cannot trade the paper's market. Yang studies KOSPI 200 names using Korean-language Naver Finance articles. Our tradable universe would consist of liquid US-listed equities, with timestamped US news and press releases. Our English tone classifier was trained and evaluated only on historical text.
Reproduction is out because we lack the Naver corpus, the KOSPI 200 constituent history, the Korean firm characteristics and KrFinBERT's Korean tone labels. Any US run of ours therefore evaluates an English-text substitute. It does not evaluate Yang's result, and none of its numbers belongs beside his 0.242% as if both runs measured the same strategy.
Thirteen months of Naver Finance
The proposed mechanism is short-horizon cross-sectional underreaction to news tone. The strategy scrapes firm-level news, scores positivity, buys the names with the highest tone yesterday and shorts those with the lowest. It holds for one day and repeats.
Yang scraped Naver Finance for the 200 constituents of the KOSPI 200 from May 2023 to May 2024, collecting 502,737 articles. Three filters reduced the set to 497,217. The company name had to appear within the first 25 words, receive at least twice a mention, and the article had to contain at least 30 words. That works out to roughly 2,486 articles per firm.
Tone is supplied by five open-weight models and KOSELF, the Korean finance sentiment lexicon used as the dictionary benchmark. Four of the five models are fine-tuned for Korean, while DistilBERT is multilingual. Only two of those four also received fine-tuning on financial text. For the encoder models, the score for each sentence is positive minus negative, divided by positive plus negative, then averaged across sentences. Llama3 receives a prompt requesting a number in [-1, +1]. The consensus tone for a firm on day t-1 is the mean across its articles that day. The main tables use equal-weighted quintile and tercile sorts, rebalanced daily.
KrFinBERT, the Korean finance-tuned BERT, predicts next-day return above the KOSPI with a coefficient of 0.257 (t=3.207). Its equal-weighted quintile high-minus-low portfolio earns 0.242% per day (t=3.772), which Yang annotates as 5.324% per month. Factor adjustment barely moves the result. The Fama-French three-factor alpha is 0.232% (t=3.582), and the five-factor alpha is 0.232% (t=3.566).
Across the thirteen months, the long-short portfolio compounds past 60%, versus roughly 10% for the KOSPI. The t-stats are significant, although the regressions account for only 0.2 to 0.3 percent of the variance.
Why does the 2018 encoder win?
When all five tone measures and KOSELF enter together, BERT alone survives at 0.228 (t=2.657). RoBERTa changes sign to -0.062 (t=-0.401), while KOSELF falls to 0.269 with a t of 0.718. The portfolio tables tell a less flattering story for the language models. KOSELF's dictionary high-minus-low earns 0.151% per day (t=3.016), ahead of RoBERTa's 0.146%, ELECTRA's 0.134%, Llama3's 0.060% and DistilBERT's insignificant 0.065% (t=1.388).
Four of the five language models add nothing over a Korean finance word list.
The winner is a 110M-parameter encoder from 2018, fine-tuned on Korean and financial text. Yang credits this dual adaptation rather than the architecture. BERT and RoBERTa are the only two models with both adaptations, and they rank as the two strongest performers.
Llama3, the generative model, has the lowest point estimate among the five at 0.060%, despite statistical significance (t=2.336). Value weighting pushes it to -0.006% (t=-0.139). Adapting an inexpensive encoder to the language and domain pays better here than spending effort on the newest decoder.
Smaller members carry the result
Value weighting lowers BERT from 0.242% to 0.168% (t=2.513). It removes significance from the three other models that previously had it, while DistilBERT never did. Yang states: "only BERT retains statistical significance, while all other models lose significance." The return is concentrated among the smaller members of a large-cap index.
KrFinBERT's double sorts provide less separation than the abstract's claims about lower size, coverage and foreign ownership imply. Small firms deliver 0.221% (t=2.639), compared with 0.200% (t=2.345) for large firms. Low analyst coverage produces 0.216% (t=2.162), versus 0.211% (t=3.057) for high coverage. Growth firms earn 0.221% (t=2.434), against 0.178% (t=2.507) for value firms. We did not find a test comparing the subgroup spreads.
Foreign ownership shows the clearest separation, again in the BERT rows. Low ownership returns 0.336% (t=3.288), while high ownership returns 0.143% (t=2.123).
Most of the performance comes from the long leg, at about +50% cumulative, against about -10% for the short leg. This distinction carries extra weight because Korea imposed a full short-selling ban on listed stocks from November 2023 through the end of the sample. We found no mention of the ban anywhere in the paper. The 2024 subperiod result of 0.158% per day (t=2.357), measured over five months, falls entirely within that ban.
Four unresolved implementation choices
The filters, scoring formula, sort, rebalance frequency and HuggingFace model identities are specified. Four things remain open:
- No intraday cutoff is stated. Day t-1 may include after-hours releases, leaving tomorrow's tradability dependent on an unspecified timestamp rule.
- The next-day return convention is unspecified. Close-to-close and open-to-close can produce very different results for a news signal, and they change the cost arithmetic.
- Firms with no articles on day t-1 appear to leave the universe. Daily breadth then follows news flow rather than a fixed count.
- Llama3's identity is unclear. Table 2 reports 3B parameters beside the model name llama3-70b-8192.
Costs enter as a flat subtraction rather than through a turnover-weighted net-return series. Presumably for this reason, the conclusion still names transaction costs and liquidity constraints among the issues the study does not assess. The paper reports no Sharpe, drawdown or hit rate.
We have written before about a paper that specifies its estimation precisely while saying little about the version an investor would trade (our EGARCH note). This paper is more candid about the gap, yet the net series remains unbuilt.
Show the long leg after costs
The result that would change the assessment is a net-of-cost daily return series for the long leg alone, covering a period that includes 2022 and 2025 and using a specified news cutoff. The long leg generated the 50% cumulative return and requires no borrow.
Its cost advantage is smaller than its contribution to profit might suggest. A long-only portfolio still pays the 0.20% sell tax whenever it rebalances, plus roughly half of the 0.03% round-trip brokerage. Yang's subtraction leaves 0.012% from the 0.242% gross spread. Using his two figures, our arithmetic puts explicit costs alone at 95% of gross.
Our own US run measures an English-text substitute in a different universe. It cannot assess this Korean strategy.