UnrestSent200K

UnrestSent200K · A Bangla crisis-sentiment benchmark · 199,399 human-labelled comments · July–August 2024

Can LLMs follow the pulse of a crisis?

Evaluating crisis sentiment in Bangladesh's July uprising. A 200K-comment Bangla benchmark in which every comment keeps its parent post and its moment in the event, so models can be tested on context and on time, not only on words.

Md. Samiul Alim*Mahir Shahriar Tamim*Tanvir Ahmed KhanSharjil KhanRafia Ferdous DutiShahriyar Zaman RidoyMohammad Ali Moni

North South University, Dhaka, Bangladesh · Charles Sturt University, Australia
* equal contribution · samiul.alim01@northsouth.edu

Comments per day across the five phases, drawn as a pulse
Comments per day · Jul 5 → Aug 31, 2024 · five phases
peak: 7.7K/day · Aug 5–10
199,399human-labelled Bangla comments
5event-aligned phases in 58 days
2,000+source posts · Facebook & YouTube
0.73Cohen's κ · 14 native annotators
94.2%blind-audit agreement
21models benchmarked

The event

Seven weeks, five phases, one population

Between 5 July and 31 August 2024, public conversation in Bangladesh moved through early mobilisation, a nationwide internet blackout, the collapse of the government, a political transition and, finally, a flood crisis. The language, the platforms and the people posting stayed largely the same. The context changed almost overnight.

That is the stress test this benchmark is built around. A short comment that reads as praise on one day can be sarcasm the next, and its meaning often depends on the news post it answers. Most sentiment benchmarks assume language is stable, judge each utterance alone, and are built for English. UnrestSent200K is built to break those assumptions on purpose.

Every comment in the dataset is timestamped and assigned to one of five phases that follow the event itself. The phases are used for temporal analysis and split construction; they are never used as labels. The chart shows how the volume of public comment surged after the blackout ended and the regime fell, then shifted again as floods became the dominant concern.

Two questions follow from this structure. Can a model trained during one phase still read sentiment in the next? And can it read a comment at all without the post that provoked it?

Volume and sentiment across the five phases

Hover or focus a phase for counts · 199,399 comments

  • Negative
  • Neutral
  • Positive
Table view
PhasePeriodDaysNegativeNeutralPositiveTotalPer day
P1 · Jul 5–15Pre-Escalationearly mobilisation · 6,727 comments
P2 · Jul 16–Aug 4Crisis & Blackoutnationwide internet blackout · 33,684
P3 · Aug 5–10Post-Blackoutregime collapse · 46,197
P4 · Aug 11–20Post-Revolutionpolitical transition · 69,536
P5 · Aug 21–31Flood Crisisoverlapping flood crisis · 43,255

The gap

Existing Bangla resources are static, isolated and consumer-oriented

SentNoB, SentiGOLD and BnSentMix made real progress, but none divides a single unfolding event into phases, links comments to their parent posts, or supports controlled comparisons across time.

Bangla social-media benchmarks by size

Comments or posts, thousands · sentiment, hate-speech and offensive-language datasets

UnrestSent200K is 2.8× the size of SentiGOLD, the largest previous Bangla sentiment corpus, and the only one anchored to a single crisis.

Static

Benchmarks treat language as stable over time. A crisis moves faster than retraining cycles, and sentiment can pivot around a single event.

Decontextualised

Each utterance is judged in isolation. In a comment thread the post is often part of the meaning, not extra information.

Consumer-oriented

Existing Bangla corpora cover reviews and general social text. Politically sensitive crisis discourse, with sarcasm and implicit references, is different.

The dataset

Every comment keeps its post and its moment

Public posts and comment threads were collected from verified news pages, public-interest groups and high-engagement discussion threads on Facebook and YouTube, using Apify, the Facebook Graph API and the YouTube Data API v3.

Download

Get UnrestSent200K

One archive with the full labelled corpus, the parent-post headlines, and the code used for the encoder and LLM experiments in the paper.

Released for research use. Comments are de-identified: usernames, profile links and other identifying information were removed before release. Please cite the paper if you use the data.

  • Inside the archive

  • UnrestSent200K.csv199,359 labelled rows · columns: combined_text, label, post_date · 102 MB unpacked
  • July_Post_And_Label.csv2,514 parent posts · columns: date, Headline
  • Encoder_Fine_Tune.ipynbNotebook for fine-tuning the encoder baselines
  • Zero_and_FewShot_LLMS.pyZero-shot and few-shot prompting script for the LLM benchmark

What each record carries

  • CommentBangla text, NFC-normalised; emojis kept because they carry sentiment
  • Parent postthe news post or headline the comment answers
  • Timestamp & phaseone of the five event-aligned phases
  • PlatformFacebook or YouTube
  • Sentiment labelpositive, neutral or negative, assigned to the comment only
  • Confidenceannotator confidence in [0, 1], mean 0.82

Class balance

  • 47.6% negative · 94,830
  • 21.0% neutral · 41,784
  • 31.5% positive · 62,785

Comments shorter than three tokens, mostly non-Bangla text and obvious spam were removed. Near-duplicates within a thread were dropped at a Jaccard threshold of 0.85. URLs and user mentions were replaced with placeholder tokens and all identifying information was stripped. Splits are 70/15/15, stratified jointly on phase and sentiment so every partition keeps the same class and temporal mix.

Overview diagram of the UnrestSent200K framework: data acquisition and contextual grounding, the PiLA annotation pipeline, and the experimental probing axes for temporal robustness and contextual grounding
Construction and evaluation framework. (1) Comments are linked to their parent posts and metadata, producing utterance-only and context-aware inputs. (2) PiLA, a five-stage fully human annotation workflow. (3) Two probing axes: temporal robustness across phases, and contextual grounding across encoder and LLM families.

Stratified split by phase and sentiment

Train 70% · validation 15% · test 15% · counts

PhasePeriodTrain negTrain neuTrain posVal negVal neuVal posTest negTest neuTest posTotal
P1 Pre-EscalationJul 5–152,2818531,5764891833384881823376,727
P2 Crisis & BlackoutJul 16–Aug 49,1804,7009,6981,9671,0072,0781,9681,0082,07833,684
P3 Post-BlackoutAug 5–1014,1537,48210,7043,0331,6032,2943,0321,6032,29346,197
P4 Post-RevolutionAug 11–2023,66410,53214,4795,0712,2573,1035,0702,2573,10369,536
P5 Flood CrisisAug 21–3117,1045,6827,4933,6651,2181,6063,6651,2171,60543,255
Overall66,38229,24943,95014,2256,2689,41914,2236,2679,416199,399
Stacked bar chart of sentiment share by comment length: short comments are half positive, and the neutral share rises steadily with length while negative stays near 30 percent
Sentiment by comment length. Short comments (1–10 words) skew positive; the neutral share rises with length as comments turn to explanation and fact. Negative stays flat.

Post-level framing

Besides the comment label, each source post carries a coarse framing label, Hope, Despair or Outrage, that describes the dominant frame of the post. These labels were used as annotation context and are released as metadata. The prediction task stays comment-level.

Why both input settings are released

The corpus supports comment-only and post + comment evaluation side by side. In crisis discourse, removing the post can remove the information a correct label depends on, so the dataset lets future work measure how well a model actually uses discourse context.

Annotation · PiLA

Five stages, every label by a person

Pilot Label Annotation (PiLA) is a fully human protocol. Fourteen native Bangla-speaking undergraduates with first-hand exposure to the events labelled every comment; ten senior validators, drawn from a separate pool, adjudicated and audited. The whole process took about eight months.

1

Guideline pilot

Ambiguity classes identified and documented: sarcasm, rhetorical questions, religious expressions, political references, code-switched profanity, emoji sentiment. Guidelines refined through pilot annotation, then frozen.

2

Training & qualification

Annotators trained on gold examples and admitted only at Cohen's κ ≥ 0.65. All 14 qualified, scoring 0.66 to 0.81.

3

Context-aware double annotation

Two annotators label each comment independently, with the parent post as context but never copying its sentiment. Gold-check probes inserted at 2%; drift triggers retraining.

4

Senior adjudication

The 21.6% of comments with disagreement go to a senior validator; 3.1% escalate to a three-validator panel. Both labels, confidences and the cited rule are logged.

5

Independent blind audit

Two validators not involved in adjudication relabel 2,000 stratified comments, 400 per phase, blind to gold: 94.2% exact agreement, 96.8% within one step.

0.73pairwise Cohen's κ · 0.69–0.78 by phase
0.71Krippendorff's α · full pool
78.4%direct agreement before adjudication
0.82mean annotator confidence · 87.3% ≥ 0.7
94.2%blind-audit exact agreement
8 mototal annotation effort
Diagram of the five PiLA stages from guideline pilot through training, double annotation, senior adjudication and blind audit, ending in the reliability summary
The PiLA pipeline. Guideline pilot, qualification, double annotation, adjudication and blind audit, ending in a labelled corpus with logged reliability.
Two bar charts of the annotator and validator pools by academic background and gender
Who labelled the data. Annotators (n = 14): 71% computer science, 29% electrical engineering; 57% male, 43% female. Validators (n = 10): 60% / 40% on both. All native Bangla speakers residing in Bangladesh; informed consent obtained and the protocol reviewed by the host institution.

Result 1 · Context

The post is part of the meaning

A paired ablation with everything held fixed except the input: four fine-tuned encoders see either the comment alone or the parent post and comment joined by a separator token. Every architecture gains 7 to 11 accuracy points.

Accuracy with and without the parent post

Random 70/15/15 split · same corpus, optimiser, schedule and hyperparameters

  • Comment only
  • Post + comment
Table view
ModelAcc, commentF1, commentAcc, post + commentF1, post + commentGain (pp)

The gain is similar across very different architectures, which makes it a property of the task rather than of any single model. Comment-only sentiment classification in a crisis thread systematically measures the wrong thing.

The best encoder result combines Bangla pretraining and context: grounded BanglaBERT at 74.72% accuracy and 0.717 macro-F1, with balanced precision and recall.

Same comment, two readings

Comment aloneকারা করছে, আমরা জানি। স্বাধীন দেশের মানুষেরা।“We know who is doing this. People of an independent country.”
Neutral
With its parent post
স্ত্রী ও পুত্রকে নিয়ে এক কাপড়ে বাড়ি থেকে বের করে দেওয়া হয় রাহুল আনন্দকে“Rahul Anand was forced out of his house with his wife and son, with only the clothes they were wearing.”
Negative

Alone, the comment reads as a flat statement. Under the headline it becomes bitter irony about an “independent” country. Gold label: negative.

Red spray-painted graffiti on a wall in Dhaka reading 5th Aug, 2nd Victory
“5th Aug · 2nd Victory.” Graffiti from the July 2024 uprising. The day the regime fell is also the day the corpus pivots: the largest sentiment shift in the data sits at the boundary between the blackout and what came after.

Result 2 · Time

Train in July, fail in August

Encoders fine-tuned on comments from 5 July to 5 August were tested on two later windows that do not overlap the training period: Test-A (6–20 August) and Test-B (21–31 August). Same language, same platforms, same event. Only the stage of the crisis changed.

Every model pays for the gap. BanglaBERT stays the strongest and degrades least, reaching 54.83% on Test-B, still twenty points below its in-distribution result with the same post + comment input. Across the four encoders the loss is 19 to 28 macro-F1 points.

The lexical distribution itself moves: Jensen–Shannon divergence between phases sits at 0.37 to 0.38, fewer than half of the top words are shared between phase pairs, and vocabulary entropy rises from 7.60 to 7.79. The drop is a property of the data, not of one model.

The implication is direct. A random test split can report a model as reliable while it fails on the same crisis a few weeks later. Crisis sentiment needs temporal splits.

Accuracy under temporal shift

Trained Jul 5–Aug 5 · tested on later windows · in-distribution result shown for reference

  • In-distribution (random split)
  • Test-A · Aug 6–20
  • Test-B · Aug 21–31
Table view
ModelIn-dist. accTest-A accTest-A F1Test-B accTest-B F1

Sentiment shift between phases, Cohen's d

All ten phase pairs · darker = larger absolute effect · hover a cell

Thresholds: negligible < 0.2 · small < 0.5 · medium < 0.8. Seven of ten contrasts are negligible; the two medium effects both touch the post-blackout phase.

Sentiment pivots. It does not drift.

If degradation over time reflected gradual drift, adjacent phases should differ by similar amounts. They do not. The matrix is sparse: the strongest changes both involve Post-Blackout, against Crisis & Blackout (d = −0.556) and against Flood Crisis (d = 0.567). The shift is concentrated at the moment internet access returned, the regime collapsed and people reacted to a new political reality, then again as the floods took over.

For crisis monitoring, smooth temporal modelling is the wrong inductive bias. This is a change-point problem, and the benchmark is built to study it.

Three small charts: Jensen–Shannon divergence between phases around 0.37 to 0.38, top-word Jaccard overlap around 0.42 to 0.47, and vocabulary entropy rising from 7.60 to 7.79
Lexical drift across phases. Divergence, vocabulary turnover and rising entropy together show measurable non-stationarity.

Result 3 · LLMs

Prompting helps. Fine-tuning wins. Neither clears 80%.

Fifteen open-source and proprietary LLMs were prompted in Bangla with an expert persona, zero-shot and with seven in-domain examples, all on the same post + comment input as the encoders. Two mid-scale models were then LoRA-fine-tuned on the training split.

78.4%Gemma-4B · LoRA SFT
best overall · κ 63.1
76.0%GPT-4o-mini · few-shot
best prompted · κ 55
74.7%BanglaBERT · post + comment
best encoder · F1 0.717
+2.4 ppGemma-4B over GPT-4o-mini
+3.7 pp over grounded BanglaBERT

Prompted LLM leaderboard

Bangla expert prompt · accuracy, macro-F1, MCC and κ in % · best per section marked

ModelAccuracyMacro-F1MCCCohen's κ

Greedy decoding at temperature 0; any response that is not exactly one of the three labels counts as an error. Few-shot examples help the strongest models most: κ rises from 47 to 55 for GPT-4o-mini and from 34 to 40 for GPT-4.1-mini, while several sub-3B open models barely move. LoRA rows use r = 64, α = 32, learning rate 3e-4.

Scale alone does not close the gap

Models at or below 1.5B parameters cluster near chance on the chance-corrected metrics. The 3 to 4B range begins to discriminate, and Gemma4-4B reaches κ 46.3 with examples. Even at 8B, LLaMA-3.1 trails the proprietary models in the same accuracy band. Fine-tuned Gemma-4B beats fine-tuned LLaMA-3.1-8B, so backbone and recipe matter more than parameter count here.

The strongest prompted model, GPT-4o-mini, lands only 1.3 points above grounded BanglaBERT and shares its failure modes. Supervised adaptation on in-domain labels remains the most reliable lever, and the 78.4% ceiling says the task is still open.

Four panels plotting accuracy, F1, MCC and Cohen's kappa against parameter size for Gemma, Qwen and Llama families, with GPT-4o-mini and BanglaBERT as reference lines and LoRA-tuned Gemma-4B as the top point in every panel
Best result per family against parameter size. Open-source few-shot models improve with scale but stay below the proprietary and fine-tuned models; LoRA SFT on Gemma-4B is the strongest point on every metric.

Three failure modes, before and after fine-tuning

The same post and comment pair, judged by Gemma-4B zero-shot and by the same backbone after LoRA fine-tuning on the training split.

Sarcasm

Postআমার বাসায় কাজ করা পিয়ন এখন ৪০০ কোটি টাকার মালিক: প্রধানমন্ত্রী“The peon who worked at my house now owns 4 billion taka: Prime Minister.”

Commentআহ কি গর্বের কথা!“Ah, what a matter of pride!”

GoldNegative
Zero-shotPositive
LoRA SFTNegative

Surface “pride” inverts under a corruption headline. Zero-shot follows the literal words; fine-tuning catches the intent.

Neutral collapse

Postদেশে ইন্টারনেট সেবায় ধীরগতি, ফেসবুক–মেসেঞ্জারে প্রবেশেও বিঘ্ন“Internet service is slow nationwide; access to Facebook and Messenger is also disrupted.”

Commentএটা হাসিনার কাজ“This is Hasina's doing.”

GoldNegative
Zero-shotNeutral
LoRA SFTNegative

Read as a factual attribution by the zero-shot model. In context it is an accusation, and the fine-tuned model learns that pragmatic reading.

Surface-cue confusion

Postবন্যায় এখন পর্যন্ত ১৩ জনের মৃত্যু, ক্ষতিগ্রস্ত ৪৫ লাখ মানুষ“13 dead in floods so far; 4.5 million affected.”

Commentআল্লাহ আপনি সবাইকে হেফাজত করেন।“May Allah protect everyone.”

GoldPositive
Zero-shotNegative
LoRA SFTPositive

Disaster vocabulary in the post drags the zero-shot label negative. The prayer is solidarity, and the fine-tuned model reads it that way.

Diagnosis

Where the models fail

Aggregate metrics say how much a model fails, not where. Topic modelling over the corpus and over the 2,789 comments BanglaBERT misclassified shows the failures are structured, and that the neutral class carries most of them.

Per-class accuracy, grounded BanglaBERT

Test split · percent of gold-labelled comments recovered

  • Negative
    83.7%
  • Positive
    63.1%
  • Neutral
    42.5%

How the 2,789 errors break down

  • 895 neutral → negative · 32.1%
  • 356 neutral → positive
  • 531 negative ↔ positive · 19.0%
  • 1,007 other

More than half of neutral test comments (57.5%) are mislabelled, with a 2.5 : 1 bias toward negative. The model is confident while wrong: mean confidence on errors is 0.977, and in one topic it reaches 1.000 while 85 neutral comments are labelled positive. Probability thresholds cannot catch these errors.

The corpus itself is hegemonic rather than pluralistic. One topic, league, government and quota, absorbs 56.3% of all documents and carries net-negative polarity, so a model that overfits its surface inherits a negativity prior on everything else. And the blackout phase is the most coherent period in the corpus (Cv 0.78 against 0.38 to 0.42 elsewhere): crisis crystallises discourse rather than fragmenting it.

Dot plot of sentiment polarity trajectories for nine topics across the five phases, with arrows showing shifts such as topic 22 on quota reform moving from plus 57 to minus 33 percent
Topic-level sentiment trajectories. Hollow circles mark initial polarity, filled circles final polarity scaled by volume. Quota reform (T22) flips from +57.1% to −33.3% across P3 → P4 (σ = 45.8), while flood-relief topics stay within σ ≈ 12–15. The largest reversals cluster at P2 → P3 and P3 → P4.

Volatility is topic-specific: six topics are highly volatile, fourteen medium, five stable. Politically charged themes show three times the sentiment volatility of humanitarian ones. Temporal robustness therefore has to be measured topic-conditionally, and recalibration should follow topic volatility profiles rather than aggregate confidence.

Findings

Five things the benchmark shows

  1. 76.0%GPT-4o-mini, few-shot

    Prompted LLMs are strong, but not sufficient.

    The best prompted model only slightly improves over the strongest post + comment encoder and shares its failure modes on sarcasm, implicit political reference and phase-dependent meaning.

  2. 78.4%Gemma-4B, LoRA SFT

    Supervised adaptation gives the best performance.

    LoRA fine-tuning on in-domain labels beats every prompted setup, and the smaller backbone beats the larger one. The ceiling is still below 80%.

  3. +7–11accuracy points from context

    Parent-post context improves prediction.

    Every encoder gains when the parent post is included. Many comments cannot be reliably interpreted in isolation.

  4. −19–28macro-F1 under time shift

    Temporal shift causes large degradation.

    Models trained on earlier phases perform far worse on later ones, though language, platforms and event are unchanged. Random splits overestimate performance in a crisis.

  5. d = 0.57post-blackout vs flood

    Sentiment changes through phase-level shifts.

    The largest changes cluster around the post-blackout period rather than accumulating gradually. Crisis sentiment is a change-point problem.

Limitations

  • Two platforms. Facebook and YouTube were central to public discussion, but X, Telegram, private messaging and offline discourse are absent.
  • Three labels. Polarity keeps large-scale annotation reliable; anger, fear, hope, grief and stance are left to future extensions.
  • One event. Results are about one crisis in one language; cross-event and cross-language studies are needed before generalising.
  • Standard models. None of the evaluated systems was designed for temporal adaptation or change-point detection. That is the open problem the benchmark is meant to serve.

Cite & resources

Use the benchmark

The release includes the de-identified corpus, fixed splits, parent-post context, timestamps, platform and phase metadata, annotation guidelines, prompt templates, preprocessing and evaluation scripts, and the encoder and LoRA configurations.

BibTeX

@misc{alim2026pulse,
  title   = {Can LLMs Follow the Pulse of a Crisis? Evaluating Crisis Sentiment in Bangladesh's July Uprising},
  author  = {Alim, Md. Samiul and Tamim, Mahir Shahriar and Khan, Tanvir Ahmed and Khan, Sharjil and Duti, Rafia Ferdous and Ridoy, Shahriyar Zaman and Moni, Mohammad Ali},
  year    = {2026},
  note    = {UnrestSent200K: a Bangla crisis-sentiment benchmark. Dataset and code: https://github.com/sami0055/UNRESTSENT200K}
}

Ethics

All data were collected from publicly accessible Facebook and YouTube content in accordance with platform terms of service. Usernames, profile links and other identifying information were removed during preprocessing. Annotators gave informed consent and the protocol was reviewed by the host institution. LLMs were used as experimental systems and, in a limited way, for language editing and figure assistance; all such material was reviewed by the authors, who retain full responsibility.