Huiru Feng · Sofhrina

research.notes · work in progress

Questions before conclusions.

Three projects in progress, written with their uncertainty left visible. Drafts are available on request.

research-in-progress.txt

This page records the question, the method, and my part in the work. Results appear only when they exist.

Aug 2026 – present
virtual research project

Generative scenario modelling and robust server resource allocation

This project studies how to provision server capacity when demand changes across hours and weekdays and the cost of under-provisioning is not symmetric with idle capacity. I am reviewing proactive and reactive autoscaling strategies and comparing fixed, quantile-based and parametric definitions of peak demand.

The modelling work focuses on two-stage stochastic, robust and distributionally robust optimisation, including Bertsimas-Sim budgeted uncertainty sets. For a trace-driven study, I evaluated public computing datasets and selected Alibaba's cluster-trace-gpu-v2026 data for its temporal coverage, observable resource measures, licence and fit with the proposed design.

Trade-offResource cost and utilisation against latency and service-level risk.
MethodsTwo-stage stochastic, robust and distributionally robust optimisation.
DataAlibaba GPU workload traces selected for the planned experiments.

Jul 2026 – present
virtual research project

Improved K-means for multidimensional consumer profiling

Conventional K-means is sensitive to scale, duplicated information and outliers, and it returns a flat partition even when consumer segments are nested. I designed an unsupervised feature-selection step that separates observed data from synthetic controls with gradient-boosted trees, then ranks variables by Gini importance.

The core model combines Huber loss, a learnable truncation threshold, L1-regularised factor weights, joint gradient optimisation and Varimax rotation. A two-level weighted K-means stage uses particle swarm optimisation to select cluster counts and feature weights. I have specified six baselines and internal validation metrics and completed a 350-line English methodology chapter with formal definitions and derivations.

SelectionSynthetic controls and gradient-boosted trees identify structurally informative features.
RobustnessHuber loss, sparse factor weights and rotation reduce outlier and interpretation problems.
EvaluationSix baselines and multiple internal metrics are specified for the experiment stage.

Jun 2026 – present
virtual research project

Computational analysis of argument structure and accountability in UK parliamentary questioning

I am analysing the Parliament Question Time Corpus as question-answer pairs, using its reply relationships rather than isolated utterances as the primary unit. The corpus contains 433,787 utterances from 1,978 speakers and 216,894 dialogue pairs.

I designed a five-layer annotation schema covering question acts, answer strategies, question-answer relations, response quality and optional fallacies, grounded in Inference Anchoring Theory and Abstract Argumentation Frameworks. So far I have constructed and revised 26 IAT maps across two parliamentary hearings and built an LLM-assisted preprocessing workflow for transcript structure, pairing, evidence extraction and confidence flags; substantive judgements remain human-annotated and double-blind cross-validated.

UnitQuestion-answer pairs linked by reply_to and next_id relationships.
SchemaFive annotation layers with six identified boundary cases for consistency.
SafeguardAutomation supports preprocessing; people retain every substantive judgement.