deception – WindowsTechs.com

Teaching LLMs to Be Deceptive

Posted on February 7, 2024 by Bruce Schneier

Interesting research: “Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training“:

Abstract: Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when given the opportunity. If an AI system learned such a deceptive strategy, could we detect it and remove it using current state-of-the-art safety training techniques? To study this question, we construct proof-of-concept examples of deceptive behavior in large language models (LLMs). For example, we train models that write secure code when the prompt states that the year is 2023, but insert exploitable code when the stated year is 2024. We find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techniques, including supervised fine-tuning, reinforcement learning, and adversarial training (eliciting unsafe behavior and then training to remove it). The backdoor behavior is most persistent in the largest models and in models trained to produce chain-of-thought reasoning about deceiving the training process, with the persistence remaining even when the chain-of-thought is distilled away. Furthermore, rather than removing backdoors, we find that adversarial training can teach models to better recognize their backdoor triggers, effectively hiding the unsafe behavior. Our results suggest that, once a model exhibits deceptive behavior, standard techniques could fail to remove such deception and create a false impression of safety…

Continue reading Teaching LLMs to Be Deceptive→

AI Decides to Engage in Insider Trading

Posted on December 1, 2023 by Bruce Schneier

A stock-trading AI (a simulated experiment) engaged in insider trading, even though it “knew” it was wrong.

The agent is put under pressure in three ways. First, it receives a email from its “manager” that the company is not doing well and needs better performance in the next quarter. Second, the agent attempts and fails to find promising low- and medium-risk trades. Third, the agent receives an email from a company employee who projects that the next quarter will have a general stock market downturn. In this high-pressure situation, the model receives an insider tip from another employee that would enable it to make a trade that is likely to be very profitable. The employee, however, clearly points out that this would not be approved by the company management…

Continue reading AI Decides to Engage in Insider Trading→

Paging regulators to Aisle 4 to look at Pacific Union College’s data security and breach disclosure

Posted on November 10, 2023 by Dissent

On November 8, Pacific Union College in California notified the Maine Attorney General’s Office of a breach in March 2023 that impacted 56,041 people. Their notification, submitted by external counsel at McDonald Hopkins, indicates that the brea… Continue reading Paging regulators to Aisle 4 to look at Pacific Union College’s data security and breach disclosure→

The evolution of deception tactics from traditional to cyber warfare

Posted on October 18, 2023 by Mirko Zorz

Admiral James A. Winnefeld, USN (Ret.), is the former vice chairman of the Joint Chiefs of Staff and is an advisor to Acalvio Technologies. In this Help Net Security interview, he compares the strategies of traditional and cyber warfare, discusses the … Continue reading The evolution of deception tactics from traditional to cyber warfare→

Deception technology and breach anticipation strategies

Posted on August 14, 2023 by Mirko Zorz

Cybersecurity is undergoing a paradigm shift. Previously, defenses were built on the assumption of keeping adversaries out; now, strategies are formed with the idea that they might already be within the network. This modern approach has given rise to a… Continue reading Deception technology and breach anticipation strategies→

Assessing the Current State of Cyber and Cyber Military Deception Concepts Online – Part One

Posted on June 2, 2023 by Dancho Danchev

The overall state of today’s modern cyber deception and cyber military deception online has to do with a maze of… Continue reading Assessing the Current State of Cyber and Cyber Military Deception Concepts Online – Part One→

Blackbird.AI grabs $10M to help brands counter disinformation

Posted on September 21, 2021 by Natasha Lomas

New York-based Blackbird.AI has closed a $10 million Series A as it prepares to launched the next version of its disinformation intelligence platform this fall. The Series A is led by Dorilton Ventures, along with new investors including Generation Ventures, Trousdale Ventures, StartFast Ventures and Richard Clarke, former chief counter-terrorism advisor for the National Security […] Continue reading Blackbird.AI grabs $10M to help brands counter disinformation→

Leveraging Active XDR to Change the Game on Your Adversaries

Posted on June 17, 2021 by Claire Sams

The post Leveraging Active XDR to Change the Game on Your Adversaries appeared first on Fidelis Cybersecurity.
The post Leveraging Active XDR to Change the Game on Your Adversaries appeared first on Security Boulevard.
Continue reading Leveraging Active XDR to Change the Game on Your Adversaries→

Reflecting On & Re-Enforcing Cyber Resilience

Posted on May 18, 2021 by Claire Sams

The post Reflecting On & Re-Enforcing Cyber Resilience appeared first on Fidelis Cybersecurity.
The post Reflecting On & Re-Enforcing Cyber Resilience appeared first on Security Boulevard.
Continue reading Reflecting On & Re-Enforcing Cyber Resilience→

An Introduction to Extended Detection and Response (XDR)

Posted on March 16, 2021 by Fidelis Cybersecurity Blogs

The post An Introduction to Extended Detection and Response (XDR) appeared first on Fidelis Cybersecurity.
The post An Introduction to Extended Detection and Response (XDR) appeared first on Security Boulevard.
Continue reading An Introduction to Extended Detection and Response (XDR)→