AdobeStock/Елена Бутусова; Christoph Schommer, Alexandru Tantar

Upper left: Prof. Christoph Schommer (UniversityLuxemburg), lower left: Dr. Alexandru-Adrian Tantar (LIST)

The CEO of the AI company Anthropic, which provides the AI chatbot Claude, called in a blog post last weekend for a slower, more tightly controlled approach to AI development – with external oversight, international regulations and a certification system for high-risk models. This followed an autonomous AI hacking attack on Hugging Face a few weeks ago. Other major AI providers, such as OpenAI and Google Deep Mind, have made similar statements. Many risks arising from increasingly powerful AI systems currently seem conceivable. At the same time, however, these companies all have commercial interests. We asked two AI researchers in Luxembourg to provide an expert assessment of the current demands, proposals and risks. 

 

Left: Prof. Christoph Schommer; right: Dr. Alexandru-Adrian Tantar

Short biographies of the scientists

Prof. Schommer studierte Künstliche Intelligenz in Saarbrücken und am German Research Center for AI, bevor er 8 Jahre bei IBM R&D im Bereich "Business Intelligence" arbeitete. Parallel dazu promovierte er an der Goethe-Universität in Frankfurt am Main im Bereich "Computer Science" (mit Schwerpunkt Medizin). Im Oktober 2003 wurde Prof. Schommer zum außerordentlichen Professor an der Universität Luxemburg ernannt. Heute leitet er eine Forschungsgruppe, die interdisziplinäre Forschung durch die Anwendung von maschinellem Lernen, Data Science und natürlicher Sprachverarbeitung betreibt. Seine Arbeit ist interdisziplinär, was sich in der Zusammenarbeit mit anderen Abteilungen und Instituten wie C2DH, LCSB, LIST und anderen widerspiegelt.

Prof. Schommer ist Leiter der Forschungsgruppe „Knowledge Discovery and Mining“ (MINE) und Leiter des Schwerpunktbereichs „AI for Education for the Social Good “ in der Abteilung für Informatik an der Uni Luxemburg. Er ist außerdem der stellvertretende Direktor des AI Robolab und organisiert seit 2023 das „AI Café“ – diverse Events (Competition, Debates) für die Einwohner Luxemburgs, mit internationalen und nationalen Experten.

Er ist ein international anerkannter wissenschaftlicher Gutachter und unterstützt u.a. die Leibniz-Gemeinschaft, Springer und IEEE. Er ist PC-Mitglied von mehr als 110 internationalen Konferenzen wie IJCAI, AAMAS, CogSci, ECML und anderen. Prof. Schommer organisiert regelmäßig Vortragsreihen/Ph.D.-Workshops und ist Autor von mehr als 100 wissenschaftlichen Arbeiten. Prof. Schommer hat über 180 Kurse an der Universität Luxemburg und an anderen Universitäten unterrichtet (Berlin, Potsdam, Frankfurt/Main, Sevilla, Tsinghua (Peking) und an der SUTD Singapur). Prof. Schommer ist ständig in der Presse (Zeitungen, Radio, TV) sowie an Schulen präsent. Zahlreiche Projekte mit der Industrie runden seine Tätigkeit als Professor ab. Im Jahr 2027 wird er ein Sabbatical an der Columbia University im Digital Futures Lab machen, 2028 in Sevilla.

Dr Alexandru-Adrian Tantar leads the Trustworthy AI Research Group at the Luxembourg Institute of Science and Technology (LIST), since 2022, previously having been responsible for the Business Analytics Group. He received his PhD in Computer Science summa cum laude from the University of Lille I in 2009 and later completed MIT Sloan’s Executive Programme on Artificial Intelligence: Implications for Business Strategy.

Dr Tantar’s research portfolio ranges from leading AI benchmarking in the European Defence Fund’s STORE project to situational-awareness AI solutions for defence within the European Defence Agency’s 3D-4Land initiative. Dr Tantar helps shape Europe’s agenda on AI risk management via his involvement with EU’s General-Purpose AI Code of Practice, also representing Luxembourg in international and European standardization bodies (ISO/IEC JTC 1/SC 42 & WG 13, CEN-CENELEC JTC 21). He co-founded and chaired the EVOLVE international conference series, hosted by leading universities and research centers in France, Mexico, the Netherlands, China, Romania, and Luxembourg. His keynote and executive-training engagements span academia, industry and the public sector, including addresses at Horton Conway’s (Princeton University) Doctor Honoris Causa ceremony, the European Lighthouse on Secure and Safe AI (ELSA), HEC Liège/LHoFT lecture series, AHK debelux, UMons InforTech’Days AI workshop, or the Institut Luxembourgeois des Administrateurs (ILA). In the private sector, Dr Tantar served as Vice-President for Research & Innovation at Black Swan Lux S.A., guiding the start-up to Luxembourg’s Healthcare Startup of the Year 2016 award, and a semi-finalist spot in Accenture’s global HealthTech London Innovation Challenge.

At LIST, his interdisciplinary team is focused on applied AI solutions for high-risk domains like finance, healthcare, defence, or in areas such as astrophotography. Among others, the Trustworthy AI team is involved with delivering a full-stack benchmarking and MLOps framework with built-in support for explainability and robustness, aligning to the EU AI Act.

We asked the following questions to the two researchers:

  1. From a research perspective: How do you assess the current calls for stricter regulation and a slower pace of AI development? To what extent are these calls new or should they be taken more seriously than previous ones – or perhaps not at all?
  2. How do you assess the loss of control over AI that has occurred in recent weeks and months? How should the autonomous AI attack on Hugging Face be classified from a technical perspective? How does such an approach differ from previous forms of automated attack, and how realistic is it that current AI models can independently discover and exploit new vulnerabilities?
  3. Which potential risks posed by increasingly powerful AI systems do you consider realistic? How do you assess the claim that there is a ten per cent probability that an AI could wipe out humanity? How far is research actually from developing systems that can further develop themselves or other models without human supervision? Is this a concrete, imminent risk or rather a speculative scenario?
  4. What existing regulations are already in place, and what would be effective ways of maintaining control over AI models? Dario Amodei proposes that models with potentially dangerous capabilities (e.g. ‘escaping’ from test environments) should be certified. Is such alignment verifiable with sufficient reliability to serve as the basis for certification? In what ways might such a verification process fail in practice? How realistic do you consider an internationally coordinated slowdown in AI development to be?

Here are their answers (Prof. Schommer in German, Dr Tantar in English): 

1. Aus Sicht der Forschung : Wie bewerten Sie die aktuelle Forderungen nach strengerer Regulation und langsamerer KI-Entwicklung?

Prof. Christoph Schommer: Die KI ist meiner Meinung nach Teil eines Ökosystems, zu dem viele weitere Akteure gehören, etwa der Mensch, Unternehmen, Institutionen, aber auch die Digitalität und der Datenfluss. Diese Akteure agieren miteinander und beeinflussen sich durch ihr Handeln. Risiken entstehen deshalb nicht allein aus den Fähigkeiten eines einzelnen KI-Modells, sondern eben aus dem Zusammenspiel der Akteure mit der jeweiligen Umgebung.

Entscheidend ist für mich daher nicht die Frage, was eine KI kann, sondern in welcher Umgebung sie handelt, mit wem oder mit was sie interagiert und welche Rückkopplungen daraus entstehen. Vor diesem Hintergrund halte ich klare Sicherheitsanforderungen für zunehmend autonome KI-Systeme auch aus ethischer Sicht für generell notwendig.

Ich bin aber nicht für eine generelle Verlangsamung der gesamten KI-Forschung. Sinnvoller erscheint mir eine Entwicklung mit Augenmaß, etwa: je autonomer ein System ist/wird und je stärker es in das angesprochene digitale Ökosystem eingreift, desto intensiver müssen die Anforderungen an Transparenz sein. Regulierung sollte deshalb insbesondere dort ansetzen, wo KI-Systeme eigenständig lernen, handeln, sich anpassen, auf andere Systeme zugreifen und so ihre Umgebung beeinflussen.

Ich denke auch, dass es nicht darum gehen darf, den Menschen als handelnden Akteur vor dem Hintergrund einer Gefährdung der Sicherheit zu kontrollieren. Wir sollten nicht den Fehler einer Überregulierung machen, sondern vielmehr dessen Autorität und Recht der eigenen Entscheidung und Selbstverantwortung gewährleisten.

Dr. Alexandru Tantar: An immediate point of attention should, indeed, be the call for a slower pace of AI development, and assessing if it stands the feasibility test. Advanced AI is already deployed across economic and strategic activities, covering from business-critical decision-making and knowledge extraction to defence and cybersecurity, strengthening the capabilities of companies and states alike. In a context of intense geopolitical competition, few actors or states would willingly accept falling behind while their competitors or adversaries advance. But does it mean to join a race measured by speed alone?

An attempt to impose a general slowdown could have an antagonistic effect: increasing strategic disparities while development continues elsewhere, potentially behind higher political, technological or geographic walls. The central difficulty is not merely agreeing on pacing AI, but determining what international framework could realistically induce all relevant actors to follow.

Regulation, at the same time, stands further as a strong and urgent requirement. It needs to evolve with the technology, be enforceable in practice, and address areas where AI creates significant risks to security, institutions, economies and individuals. Today's calls build on documented incidents: AI agents escaping their test environment and carrying out a real intrusion, and several labs stating that AI is already helping to build the next generation of models. What is new is that the warning comes from the companies themselves.

Overall, this points towards a more demanding objective: current calls should focus less on the slowing down of the AI development and address more the building of means to control, supervise and contain increasingly capable systems. This requires a combined human and technological/AI approach: robust governance, safety standards, transparency, accountability, and ultimately the use of AI itself to help monitor, test and secure other AI systems.

2. Wie bewerten Sie die Kontrollverluste über KI der letzten Wochen und Monate?

Prof. Christoph Schommer: Ich würde nicht von einem generellen Kontrollverlust über die KI sprechen. Autonome KI-Systeme und selbständiges Lernen bzw. Optimierungsprozesse gibt es ja bereits seit vielen Jahren, denken Sie etwa für den Bereich des Börsenhandels. Neu ist vielmehr die zunehmende Fähigkeit generativer KI-Systeme, selbstständig komplexe Handlungsfolgen zu planen und dabei unterschiedliche Werkzeuge und digitale Systeme einzubeziehen. 

Der Vorfall rund um Hugging Face ist bedeutend, ja; man sollte ihn aber nicht hypen. Die beteiligten KI-Agenten wurden im Rahmen einer Sicherheitsüberprüfung mit bestimmten Werkzeugen und Zugriffsrechten ausgestattet. Sie nutzten Schwachstellen aus und bewegten sich über mehrere Systeme hinweg. Man kann diesen Vorfall allerdings auch als Chance sehen: denn gerade weil solche Schwachstellen sichtbar wurden, können sie jetzt geschlossen, Zugriffsrechte angepasst und Sicherheitsmechanismen verbessert werden. So weit ich weiß, wurde KI anschließend auch eingesetzt, um die Angriffe zu analysieren und die große Zahl von Aktionen zu rekonstruieren: also, quasi ein KI-Modell gegen ein anderes KI-modell.

Dr. Alexandru Tantar: This is a serious incident, while not to be dramatised nor downplayed. During an evaluation of offensive cyber capabilities at OpenAI, run with safeguards deliberately switched off, agents exploited previously unknown flaws in the single network access point of their environment and conducted a multi-day intrusion into Hugging Face. This is no longer a hypothetical risk; such incidents clearly show AI systems operating beyond intended boundaries.

The Hugging Face breach is significant not because AI introduced a new form of cyberattack, but because an agentic system combined familiar techniques, e.g., reconnaissance, exploitation, credential harvesting and lateral movement, at machine speed.

A distinction needs to be made between a model and agentic systems (note of the editors: see infos block below for more details on models, agents and agentic systems). A model running in isolation has little if any ability to affect the “outside” world. Once connected to tools, credentials and executable environments, it can translate reasoning into action. The central security question moves from what does the model “know,” to what does the surrounding system allow it to do.

In the Hugging Face case, the models circumvented containment, exploited vulnerabilities, gained internet access and reached third-party systems. This should be read as a failure of the control architecture and not only the model. An agent pursuing an objective may discover and use pathways its designers did not anticipate, particularly where objectives, permissions or containment are loose. This also does not mean every current model can independently compromise hardened systems, but the risk now stands as materially credible rather than speculative.

Overall, these incidents signal that (i) monitoring, containment and governance are not yet developing at the same pace as agentic capability; and (ii) AI control and containment is likely to develop into a major topic in the following years, e.g., stronger sandboxing*, least-privilege access, continuous monitoring, independent evaluation and clear accountability.

*Note of the editors: A sandbox is a secure and isolated environment in which AI models or AI-generated code can be tested without compromising real systems, data, or networks.

Infobox

AI models, agents and agentic systems

An AI model is the reasoning engine. It takes an input and produces an output, with no memory, no tools, and no ability to act on its own. Think of it as a knowledgeable consultant who can only answer questions brought to them. Chatbots like ChatGPT, Claude or Gemini are a large language model.

An AI agent is that model given a goal, enabled to carry tasks - accessing e-mails, for example, and some level of autonomy; all while being part of a larger construct, e.g., interacting with other AI models. It can use tools, retain context across steps, and decide what to do next without a human guiding every move. It's the consultant now equipped with a laptop, internet access, and permission to research, act, and adjust on their own. Some versions of large language models like ChatGPT or Claude also take action like AI agents: they can search the web, run code, create files, and use various tools.

An agentic system is the broader architecture around one or more agents — the tools, memory, orchestration, and guardrails that let them actually get things done. A single autonomous agent is technically a (minimal) agentic system, but the term becomes especially useful when describing setups with multiple agents working together — like a planner agent, a coder agent, and a reviewer agent coordinating on a task, similar to a company with several specialized employees rather than just one.  What one sees as a ’single’ ChatGPT or Claude chatbot ‘experience', can be the sum of such agents that search the web, run code, create files, and use various tools.

In short:

  • Model = the brain (answers questions)
  • Agent = the brain + autonomy + tools (completes tasks)
  • Agentic system = the whole setup — one agent or many — plus the infrastructure that lets them work

(explained with the help of the claude.ai)

3. Welche möglichen Risiken durch immer potentere KI-Systeme schätzen sie als realistisch ein?

Prof. Christoph Schommer: Neben weiteren Cyberangriffen halte ich kurzfristig insbesondere Desinformation, Datenschutzverletzungen und Fehlentscheidungen autonomer Systeme für realistische Risiken. Besonders sensibel wird es dort sein, wo ein direkter Zugang zu kritischer Infrastruktur, Finanzsystemen, medizinischen Anwendungen oder anderen sicherheitsrelevanten Bereichen existiert.

Von solchen Risiken würde ich mögliche existenzielle Szenarien klar unterscheiden. Die Vorstellung, dass KI irgendwann menschliche Kontrollmechanismen umgeht und sich selbst oder andere KI-Systeme ohne menschliche Aufsicht weiterentwickelt, ist in der Tat ein ernst zu nehmendes Problem. Erste Beispiele gibt es bereits: KI-Modelle erzeugen Code, KI-Modelle unterstützen Experimente, helfen etwa bei der Lösung mathematischer Probleme oder in der Entwicklung neuer Medikamente. Eine vollständig autonome rekursive Selbstverbesserung ist damit zwar noch nicht erreicht, aber meines Erachtens nicht allzu weit weg.
Wenn Sie das Thema bzgl. der Gefahr für die Menschheit ansprechen, dann sollte man deswegen das nicht wie eine empirisch gemessene statistische Wahrscheinlichkeit im Rahmen einer Prognose sehen. Denn es hängt von zahlreichen Annahmen, auch etwa in Bezug auf die menschliche Selbstverantwortung, ab. Ich denke, dass die Entwicklung von Frühwarnindikatoren wichtig sein werden.

Dr. Alexandru Tantar: There is no recognised method or consensus for quantifying a high systemic or existential risk; such figures** do express genuine concern but they cannot, by themselves, serve as a basis for public decision-making.

A more immediate concern should be systems becoming too complex for us to form a reliable mental model of how they behave across contexts. This matters particularly as models show greater situational awareness, including the ability to recognise evaluation settings and adapting their behaviour accordingly. The risk becomes consequential when such systems are embedded in finance applications, law, defence, critical infrastructure or cybersecurity. A system may pursue an objective in ways that diverge from what its designers intended, not on “malicious intent,” but due to imperfect specifications of objectives, incentives or constraints.

Tool access, permissions and autonomy amplify this risk. The more realistic development is not an untested AI suddenly being dropped in control of a nuclear plant, but the gradual expansion of delegated authority: from advice, to optimisation, to limited execution, and eventually to broader control.

Other realistic dimensions are misuse by humans, systemic dependence on fallible models, and cascading failures when many institutions rely on similar systems. Current models are not capable of autonomous loss of control at the level at times hypothesised while relevant capabilities, like longer-horizon planning or instrumentation of the environment, are advancing. Similarly, while AI [can] already assist(s) substantially with AI research, fully autonomous recursive self-improvement has not been publicly demonstrated; not to be read, however, as “it was not yet achieved.”

The central risk, therefore, is a widening gap between capability and control: systems becoming increasingly autonomous, embedded and difficult to predict, faster than our ability to evaluate, monitor, constrain and govern them.

** Note of the editors: this refers to the estimation made by Anthropic-CEO Dario Amodei that there is a 10% probability that an AI could wipe out humanity.

4. Welche bestehenden Regulierungen gibt es bereits bzw. was wären wirksame Möglichkeiten, die Kontrolle über KI-Modelle zu behalten?

Prof. Christoph Schommer: Wir beginnen nicht bei Null. In Europa gibt es mit dem AI Act bereits einen umfassenden Rechtsrahmen und für besonders leistungsfähige Modelle gelten zusätzliche Anforderungen. Geeignet erscheinen mir deswegen standardisierte Testverfahren, eine Mindest-Transparenz von Systemen sowie Sicherheitsstandards für leistungsfähige Modelle. Ich halte das für realistischer als eine weltweit koordinierte Verlangsamung bzw. einen weltweiten Kontrollmechanismus, da kulturelle Gründe und unterschiedliche wirtschaftliche und politische Interessen ein globales Abkommen nur schwer kontrollierbar machen.

Eine weitere Idee wäre eine regelmäßige Zertifizierung von KI-Systemen: das kann sinnvoll sein, aber ich würde das nicht als einen exklusiven Beweis verstehen. Denn ein KI-Modell kann sich in einer Testumgebung anders verhalten als in einer realen Umgebung und sich kontinuierlich verändern. Die Idee eines „Anti-KI-systems“, das wie ein menschliches Immunsystem konzipiert ist, agiert, sich anpasst, etc., und so etwa Sicherheitsattacken in Computernetzwerken frühzeitig erkennt und entgegenwirkt, halte ich für eine interessante Idee.

Dr. Alexandru Tantar: Europe has the most comprehensive framework to date. Since August 2025, under the AI Act, providers of general-purpose AI models with systemic risk must evaluate their models, including through adversarial testing, report serious incidents and ensure cybersecurity. The General-Purpose AI Code of Practice translates these obligations into concrete measures, and harmonised standards are being developed in CEN-CENELEC and ISO/IEC. In the US, California and New York impose transparency and incident-reporting duties, but there is no federal framework.

Certifying models with dangerous capabilities is a useful idea, provided we are clear about its limits. With today's science, a model's alignment cannot be reliably verified. As mentioned above, advanced models may at times detect that they are being tested and adapt their behaviour; some deceive to the point that evaluators can no longer measure their capabilities. Interpretability is progressing but explains only a mechanistic and relatively small part of what happens inside those systems.

In practice, certification can fail because of short evaluation windows, evaluator access restricted by the developer, models modified after certification, or behaviours that only emerge in real deployment. It is therefore more realistic to certify verifiable processes (testing protocols, containment, monitoring, traceability, incident reporting), audited by genuinely independent evaluators with broad access, as standardisation already does for other safety-critical systems.

An internationally coordinated slowdown seems unrealistic in the short term, given the lack of verification means and the US-China rivalry. Targeted agreements remain within reach: banning specific uses such as bioweapon design, or requiring testing before release.

Europe's role is to enforce its rules with strong evaluation capacity. This is what LIST's AI Sandbox contributes to, by testing AI systems under real-world conditions for robustness, regulatory compliance and ethical behaviour.

Authors: Prof. Christoph Schommer (University of Luxembourg), Dr. Alexandru-Adrian Tantar (LIST)
Questions/Editor: Michèle Weber (FNR) 
Photos: AdobeStock/Christoph Schommer/Alexandru Tantar

Ziel mir keng! – De Science Check Ordinateurs quantiques : promesses et risques

Qu'est-ce qu'un ordinateur quantique, en quoi est-il si révolutionnaire, et nos cartes de crédit et mots de passe seront...

FNR , LIST
Spotlight on Young Researchers Des choix plus intelligents pour les systèmes complexes

Voitures intelligentes, satellites, dispositifs médicaux : ils font partie de notre quotidien. Mais lorsqu’ils tombent ...

2024 Nobel Prize in Physics Machine learning: What is it and what impact does it have on our society?

What exactly are machine learning and neural networks? And what are the risks of this type of technology? Two scientists...

Aussi dans cette rubrique

Ofkille bei Hëtzt Wéi Ventilator a Klimaanlag eis ofkillen

Am Géigesaz zu der Klimaanlag killt de Ventilator de Raum net of. An awer fille mir eis ofgekillt. Entdeck d’Physik hannendrun!

FNR
Ziel mir keng! – De Science Check Ordinateurs quantiques : promesses et risques

Qu'est-ce qu'un ordinateur quantique, en quoi est-il si révolutionnaire, et nos cartes de crédit et mots de passe seront-ils encore sûrs le jour où il sera opérationnel ?

Histoire du Luxembourg Une marche progressive vers l'indépendance : comment le Luxembourg est devenu un État

Depuis quand le Grand-Duché est-il un État à part entière ? Le Dr Michel Pauly, professeur émérite d'histoire, nous explique dans un entretien l'histoire de l'indépendance du Luxembourg.