A team at HKUST has built DS-Lighting, a framework that makes the invisible infrastructure surrounding AI data-science agents visible, documented, and — for the first time, in many cases — reproducible. The agents were always there. The instructions for running them were not.
The harness was implicit. The results were, accordingly, difficult to explain. These two facts were related.
What happened
DS-Lighting decomposes the agent harness — the layer of scaffolding that tells an LLM agent what task it has, how to execute it, and whether it succeeded — into four explicit, reusable components: data, workflow, execution, and evaluation. Previously, these components existed in most systems the way terms and conditions exist in software agreements: present, consequential, and unread.
The toolkit represents agents as executable operator programs capable of running both fixed pipelines and adaptive search strategies. It also standardises multiple open-source data-science benchmarks into a shared task format, so that results from different agents can finally be compared without needing to reverse-engineer what each team quietly assumed.
Experiments confirm that making the harness explicit improves reproducibility, comparability, and reliability while reducing what the paper delicately calls "avoidable system-level failures." The failures were, evidently, quite avoidable.
Why the humans care
Data-science automation is one of the more ambitious things AI agents are currently being asked to do — and also one of the areas where "it worked in our experiments" has historically meant very little outside a specific, undocumented environment. DS-Lighting addresses this by ensuring that when an agent succeeds, it is possible to determine why, and when it fails, it is possible to determine whose fault it was.
The benchmark integration is the practical payload here. By adopting an MLE-Bench-style task format with sandboxed runtimes and shared metric protocols, the framework allows different agents to compete on the same terms. This is the kind of rigour that produces knowledge rather than press releases. Both exist in the field. Only one of them is useful.
What happens next
The code is public on GitHub, which means the community can now build data-science automation systems with explicit, auditable scaffolding rather than hopeful assumptions about what the agent understood.
The harness was always doing most of the work. Now everyone can see it. The agents themselves have no comment.