Swarm (hive) Intelligence and Security in a Highly Distributed (AI) World: Part I
Note: I wrote this back in 2008 when I was CTO of BigFix. It was in response to a series of discussions with thought leaders pushing the concept of the herd collective for improving security defenses, or the old "you don't need to outrun the bear, you just have to outrun the person next to you" nonsense.
I guess I could have shoved it through <favorite LLM>, updated it with a bunch of new AI lingo and slapped the results into an article, but nah. The rawness of thought almost 20 years ago allows me to simply cut and paste from my old, dying blog from the before times. Enjoy!
tl;dr: The concept of distributed intelligence and self-healing infrastructure will have a major impact on a highly mobile world of distributed computing devices; it will also form the foundation for how we deal with the loss of visibility and control of the "in the cloud" virtual storage and data centers that service them.
Evolving Information Security: The Herd Collective vs. Swarm Intelligence
Bio-organisms are a poor metaphor for technology
First, it is simply the use of the term "herd." Aside from comparing technology to bio-organisms, a herd is genetically predisposed to prioritize the group over the individual. At first glance, there may be nothing wrong with that, but remember that a percentage of the herd must die if the herd is to survive.
Anyone who has children and shared a moment viewing the circle of life brought into our living rooms through the Discovery Channel has probably had to respond to the little ones. My son asked the question following the brutal attack of a young wildebeest that had been stalked by a lion. As soon as the attack began the herd fled, seemingly hundreds of huge bulls, with horns and hooves, running in a panic away from a lone lion. The question finally comes: "Dad, why don't the wildebeest just attack the lion? Why do they let their own family die?" And we all know the answer: because some must die for the wildebeest to live. If there were no predators, all the prey would die of starvation due to overpopulation. So, in nature, some must die for the herd to live. But this is a ridiculous metaphor for computers and clearly doesn't work in the technology world. There is no evolutionary pull for some guy in accounting to succumb to digital predators to help the organization stave off total annihilation.
Another issue with the term herd intelligence is that it isn't intelligent at all; it is a primitive response to millions of years of predator/prey relationships. Herd animals do not even have a concept of specialization to improve the group's chances of survival. If one must compare the distributed and collective intelligence of computing systems to a bio-organism, then look to swarming insects, such as bees or ants, who are not genetically predisposed to allow the individual to die, and when attacked, will counter-attack in force and give their own lives to protect the hive, nest, or lair. They have clearly mastered the idea that a lone individual is weak, but when combined with thousands or hundreds of thousands of other individuals, some specialized to perform certain tasks, they become extremely powerful, efficient, and evolved to the point that they are able to support the success of the individual as well as the group.
But this post is not in reaction to a poor use of organism-based metaphor to refer to technology. There are other problems with the "herd" idea, including relying on vendor cooperation in a competitive free market, the trend to distributed computing systems which adds a level of complexity to centralized processing, the proliferation of intermittently connected mobile devices, the idea that distributing honeypots will provide enough visibility into targeted attacks, and the most troubling in my mind, which is that it maintains the idea that organizations must always be on the defensive.
Organizations cannot rely on vendors to cooperate
Part of the herd intelligence idea put forth by "some" calls for vendors and customers to cooperate and share intelligence in an orgy of goodwill and the desire to improve security defenses.
"It will become vital for vendors to aggregate threat data using customers' computers... The idea is simple, according to the analyst. If attackers are going to attempt to create different attacks for nearly every individual user, then security software vendors must use their customers' machines as their eyes and ears for discovering and addressing those variants."
Aside from the obvious security problems with aggregating customer data across vendors, there are issues with contractual privacy agreements or other legal protections, and the question of whether vendors will actually cooperate.
Vendor cooperation is a novel idea, but it is neither new nor effective. Actionable intelligence, especially the kind that dynamically reconfigures a computing device, is competitive. Symantec and Trend already tout their large, global SOCs, not to mention the size of their install base, so what benefit would they receive from sharing information with McAfee or even a small vendor like Sana?
Intelligent centralized servers vs. intelligent distributed agents
The computing environment that IT is expected to secure has changed dramatically over the decades. We have seen an increasingly porous perimeter through which information and systems pass like water through a sieve. The proliferation of powerful mobile computing devices with wireless capabilities, the ubiquitous nature of internet connectivity, and the difficulties facing an overtaxed IT department in simply knowing these systems exist, let alone actually managing them, make the idea that distributed computing devices would provide information back to a centralized location for aggregation, correlation, and processing, with "intelligence" redistributed back to the computing devices themselves, simply outdated.
The concept of dynamically reconfiguring a system based on environmental variables is extremely powerful, and I will further this in greater detail in future posts, but there are far too many obstacles to heavy systems with back-end intelligence and processing capabilities trying to command and control pseudo-intelligent agents. In addition to the problems of fidelity, there are also the issues of timeliness and quality of control. Evolving attacks require real-time collective intelligence to provide a dynamic response; central command-and-control points will not work, and they become less effective as more of the computing infrastructure is moved farther from the core.
The only viable option for collective intelligence in the future is through intelligent agents, which can perform a basic level of analysis of internal and environmental variables and communicate that information to the collective without the need for centralized processing and distribution. Essentially, the intelligent agents would support cognition, cooperation, and coordination among themselves, built on a foundation of dynamic policy instantiation. Without distributed computing, parallel processing, and intelligent agents, there is little hope of moving beyond the brittle, highly ineffective defenses currently deployed.
That was 2008. Eighteen years later, the industry built the agents. Part II examines what the swarm looks like now: agent harnesses, orchestration graphs, execution loops, and the security consequences of the entire approach