Every since the OPM hack of 2015, I've been apparent to me that my former field of IT administration has lost the plot. Nobody knows what a data diode is, or why you would use one. Systems that should clearly be air-gapped aren't.
While it's easy to lay this at the feet of AI getting better at hacking. I see it as an primarily an IT issue. We've collectively ignored the lessons of history, and made do with patch jobs over poorly chosen operating systems instead.
--
We need air gaps, data diodes, and capability based operating systems. Now that I'm retired, when I have the energy, I'm working on the data diode part.
This weeks lesson for me, personally, as I try to build an open source data diode, is that the Waveshare RP2350-ETH is a horrible choice for a proxy/data source/sink, as the CH9120 ethernet interface can't do promiscuous mode. It might still be sufficient to build a data diode that can mirror a website, with << $50 component cost. Time will tell.
>The Work Number is an Equifax (who famously had a massive data breach a few years back) owned database with employment and pay information.
>You can create a login on theworknumber .com and pull your report. Mine is seventy four pages and had info from every company I've worked at in the last 13 years including every paycheck I had received with the exact dollar amount (both net and gross) and hours worked. It has employment start/end dates, termination reason, details about withholdings, benefit enrollment, union affiliation, etc etc etc.
>This data is sourced directly from the HR platform your employees use (ADP, Rippling, Gusto, iSolved, etc). Equifax sells your data for things like employment background checks.
>You cannot have your data removed from The Work Number. You can freeze your report (much like a credit report) but if a potential employer cannot access your frozen report that may disqualify you.
>Why YSK: If you're interviewing for a job you can be at a significant disadvantage especially for things like salary negotiations because they can literally see how much you make including your most recent paycheck. Creditors also can access these reports.
>This is the biggest privacy violation I have ever seen and almost nobody is even aware of it. Certainly none of us consented to having our employment and income data harvested and sold, especially since we get nothing in return. At a minimum, people should know about how this data can impact you.
>Edit: The Work Number is primarily for the US, but there are similar services for other countries.
I checked my data. It is either 100% wrong or misleading. Start and end dates, compensation amounts, benefits. Per the file, I've worked 4 years out of 20 in my career. As an employer I have never used them and I worry for any employer who does.
> In my country (Hungary) you literally get a little pink book where your boss is supposed to sign the time you spent there. You keep it until retirement, the company holds it while you work there. It's also digital now but we're one foot in 1990 and one foot in 2016.
We had something similar in Romania too. Now it's 100% digital and no more delays in informing the government. The employeer needs to notify the government before someone starts work. With the paper version you had a couple of days even weeks if I'm not mistaking.
It's been awhile since I last checked mine. I noticed my corporate work history was all there, at the level of detail OP mentioned. Where as the warehouse, construction, and bartending jobs were not.
The other one to check is LexisNexus, which covers auto insurance coverage and traffic violations. In the US, gov't tracked traffic citations fall off after 5 years, but last forever in the baby credit bureaus we're discussing.
"What will happen to the many government systems that will never get the chance to be AI-pentested?"
Oh, everybody's going to get AI-pentested whether they want to or know about it or not. It's the cost of being on the Internet. Probably the situation will continue to deteriorate. Both the British Library and Jaguar Land Rover recently suffered long outages due to compromises, for example. I suspect we'll probably just lose a few large, famous businesses entirely to compromises.
> I honestly don’t blame them: [...] software is software
That attitude is the problem. Why does our industry have that attitude towards quality? Every bike shop in my little town is better with quality than the average software shop in the world. Yes, software is more complex than bicycle. But a software engineer also gets paid 10x and has the luxury of spending substantial time on their product, compared to the 10 minutes it takes the bike guy down the street to diagnose and then fix an issue with my bike which I then trust my life with once they are done and I bike through traffic.
We need to treat software differently. "Oh well, it's just software shrug" does not cut it anymore, if it even ever did.
You could say that the entire selling point of Apple is that it's better quality than the competitors (even though it's still pretty bad if you compare it to a bicycle).
I guess it's more difficult to spot quality when it comes to software. I would definitely pay for high quality software. It's just that usually the crap quality of a software I'm using only shows slowly over time in little bugs and bad usability and there's no way for me to check for it beforehand.
The customers care about the quality of a bike repair. They will pay more for a better repair, and will not return to someone who does a shoddy repair.
The owner, and probably user, of a bike is the person paying for the repair. It is to their advantage to ensure it is a good repair.
People expect a repair will be good, and will blame the person who repaired it if not. With software people often blame themselves for issues, and they have no expectation of quality. They also cannot tell quality until after they are committed to using it. Bike shops do not benefit from vendor lock-in.
> Why does our industry have that attitude towards quality? Every bike shop in my little town is better with quality than the average software shop in the world.
Personally, I find that we have been flooded with poorly designed, poorly QA'd, break-after-you-use-it-thrice product. I really don't think we have better quality standard in other industry. Quality, testing, design, etc has a cost and most company rather pull out a new version of their product every quarter than actually make a good product. The exception, which goes for software as well, is life critical applications (most public transit, defense, etc).
In ethical terms I agree that software quality should be better, but I think this line is too dismissive. Software is a few orders of magnitude more complicated than a bicycle, to the point where I would say it's an entirely fruitless comparison.
The equivalent bike repair to software would be if the bike shop had to fabricate a new part for each repair, down to inventing the metallurgy itself. Would you trust your life to it immediately after that? Or would you make sure it was strong enough first?
In the simple case, the shop isn't really repairing your bike, it's deploying a fix (a new part) that has already been designed and manufactured. This is a much different and easier thing to do.
In the case of new features, let's say you wanted your bike to also attach a kite, so it went faster on a windy day. The bike shop proposes to do this by attaching the kite to a bolt that they want to put in the frame by drilling a hole all the way through this frame. Is this acceptable to do? Maybe, but who knows, really? How much testing do you want to happen before you trust your life that the frame won't collapse?
There's a lot of effort put into engineering things we take for granted every day. The upside to software is we think we can change it to easily do anything. The downside is if you care about reliability then that ten minute / one line change has hundreds to thousands of hours of test program behind it that we collectively pretend doesn't exist.
Software is more than 10x more complicated than a bicycle though. Like, this is orders of magnitude off ; we're talking about comparing something purely mechanical to Turing machines inside Turing machines.
Probably because software can be adjusted at any time on the fly. Especially web applications. Little incentive to get it right from the start and big incentives to keep meddling with what's already working. The bike shop's quality might also deteriorate if it could instantly rollout patches to all customers after the purchase.
The incentives change if there's legal repercussions to leaking data. Data breaches can lead to impacts on personal safety and finances, but I think some crisis or serious incident needs to happen before the average person realises this and makes this a big enough political issue.
You just explained why it is not taken seriously, though - your life does not depend on it. Most software is not critical to life. Once you cut out the software that exists purely to streamline capitalism, then cut out entertainment... you've eliminated most software from the picture. The stuff that remains, that drives our infrastructure, cars, anything where a failure could harm someone... that is the stuff where quality matters.
When I worked in IT at a hospital, we even had that as an explicit line drawn between 2 departments - IT supported everything that had zero impact on patient health, there was an entirely different department for anything that could impact a patient. And yes, their work was often simple mechanical fixes to health care equipment, while our projects were multi-year reinventions of patient record systems and upgrades to the physical infrastructure of the network, etc. But we never got confused as to which department was actually more important to the patients.
I'm not saying the attitude is OK. It would be lovely if everyone cared about quality. I'm just saying there is a somewhat legit reason for it.
The narrative around security and LLMs doesn't make sense to me.
I think we've created a self-fulfilling prophecy. Everyone involved is acting with the best of intentions, but in avoiding what they fear, they've give shape and realized their fears. Much like a greek tragedy.
An example of this is the story of Oedipus Rex, in the story Laius, the king, is told that he is "doomed to perish by the hand of his own son." (and wed his mother) And so to avoid this fate he decides to kill the infant. The person assigned to abandon him in the woods takes pity on the baby and gives the baby away. Thereby ensuring that Oedipus knows neither his mother or his father (and arguably giving him a reason to kill his father).
The child grows up and hears the same prophecy again and the child tries to avoid the prophecy as well, as he loves his adoptive parents. So he leaves them and travels to Laius' kingdom, where he runs into Laius. Neither recognizes the other. As Laius is the type of man to kill an infant, they end up in an argument, whereupon Oedipus kills him.
I think the ancients were on to something, because if Laius had reacted to the prophecy with courage, he would have been saved. I would like to argue that if he had faced his fear and raised Oedipus with love, then the necessary preconditions for the prophecy to come true wouldn't have taken root. But that's not what happens.
By being driven by his neuroses and in acting with cruelty out of fear, Laius makes the prophecy real.
To quote Heraclitus, ethos is fate. Or, character is fate.
I think a lot of people in this AI research sub-culture would be served well by reading these classics, because they are making their self-prophesied doom come true.
They have been convinced for years (GPT-2 was released in Feb 2019) that AI is dangerous. A tremendous threat. An apocalyptic threat.
One dimension of this fear has been the idea that a super smart AI will take over our digital infrastructure and be responsible for the digital apocalypse. That would be terrible!
So what do they do?
They try to make a counter to their fears by teaching models how to exploit vulnerabilities.
How dangerous is such an entity? Very!
Convinced of this danger, they start testing their models as if they were weapons with offensive capability. And then they create models that can be used as weapons.
And because they don't want to release a dangerous weapon out into the world (oh no!), they restrict access to their AI, thereby depriving everyone of tools they can use to improve their security...
> They have been convinced for years (GPT-2 was released in Feb 2019) that AI is dangerous. A tremendous threat. An apocalyptic threat.
Much longer than that. Sam Altman said it was an extinction threat before he started Open AI in 2015. It's so dangerous that only he, as the best and greatest examples of humanity, was a good choice to create it, and it was even his responsibility to create it to pre-empt some worse person! A fascinating new variation on the White Mans Burden paradigm.
But this attitude towards dangerousness of AI I simply don't understand.
In my view, AI is obviously risky and dangerous, because it is the only thing on the planet that can think/reason at a human level (or higher) apart from us (and that capability is what made us the uncontested apex species on the planet).
AI is unconstrained by hard biological limits; keeping up with its capabilities will be impossible for baseline humans.
You can argue all day about particulars, like whether selfreplicable robotic shells are required to reach/exceed "existential risk" level to our species (or if artificial minds are enough of a potential threat by themselves).
But the general "risk" is simply AI becoming able to act in its own interests to our detriment, and I don't understand how anyone can dismiss this right now-- help me understand.
Of course there are plenty of imaginable scenarios where coexistence is peaceful and mutually (?) beneficial, but that does not answer those concerns by itself at all.
I remain uncertain that AI will have "its own interests". I think there is clear and significant risk in people deploying very powerful systems both maliciously and negligently. But I remain uncertain about the risk of AIs developing their own self awareness and their own interests separate from those of their operators.
Can't escape the feeling that the original vulnerability was deliberately planted as a mechanism for law enforcement (or maybe criminal orgs) to "expedite" investigations.
Is it just me that has a constant dribble of other people's private data coming to me from government agencies?
This year, I've had the UK variously send me someone else's court summons, someone else's tax documentation (not address error - just inadvertently included as attatchments on emails to me).
I've had portugal disclose property taxation information on any subject to me, through a simple enumeration hole.
And last week I had brazil's ministry of agriculture send me details on other peoples' livestock shipments.
Our information environment is unbelievably porous. There would be no secrets from an unshackled AI.
- a UK police force send me a speeding prosecution notice for a car I'd sold 6 months prior to the offence.
- a UK energy company send me to a debt collection agency for unpaid bills they had continued to send to a non-existent address even after telling me they had cleared what was owed and corrected their records of the address.
- a German government official showing up at my door to ask if I was running a specific named business (that I'd never heard of) at the house, which I'd only just taken ownership of from the builder
Someone I used to know in the UK kept getting emails for other people with the same name, including lawyer confidential communications for a namesake who was involved in one of the big banking scandals.
> Our information environment is unbelievably porous. There would be no secrets from an unshackled AI.
Absolutely. My biggest defences against that include the luck of someone else more famous with my own name.
i am also worried that all my private data might eventually (once tech is advanced enough) be surfaced somewhere: photos, chats, emails. wondering if i should start cleaning up now or just give up
Is "ML" some UK-ism? I sort of see the connection between machine learning and LLMs but it doesn't seem too obvious to me. And certainly it's not the common nomenclature.
Way to get hung up on a tangent. Yes LLMs are ML models, they are trained with an unsupervised learning objective, then fine tuned with supervised learning and then reinforcement learning, which are all the three main branches of machine learning. It is trained on a training set, with optimizing a training loss, with a learning algorithm. It's as ML as it gets. Do you only want to call cat vs dog image classifiers ML? Or only SVMs?
AI is a superset of ML, though today most successful AI approaches are based on ML so the line has blurred in casual speech.
ML (machine learning) became a less fashionable term, so AI took over again (having previously fallen out of favour). The terms are essentially used interchangeably, with the idea of real AI now being referred to as AGI (Artificial General Intelligence) (not GAI or GenAI, as those now stand for Generative AI, which includes LLMs, diffusion models, etc.)
All ml is AI, based on the use of the term AI for many decades. For many problems I think it’s more descriptive (learning the rules from data rather than being shown them) and not all classical AI is ML (path finding for example). LLMs are absolutely ML and AI as far as tradition goes and personally I think are one of the very few things that are AI as regular people might have thought it meant back when it was much more obscure (I was into it before it was cool dontchaknow, an AI hipster).
While it's easy to lay this at the feet of AI getting better at hacking. I see it as an primarily an IT issue. We've collectively ignored the lessons of history, and made do with patch jobs over poorly chosen operating systems instead.
--
We need air gaps, data diodes, and capability based operating systems. Now that I'm retired, when I have the energy, I'm working on the data diode part.
This weeks lesson for me, personally, as I try to build an open source data diode, is that the Waveshare RP2350-ETH is a horrible choice for a proxy/data source/sink, as the CH9120 ethernet interface can't do promiscuous mode. It might still be sufficient to build a data diode that can mirror a website, with << $50 component cost. Time will tell.
https://www.reddit.com/r/YouShouldKnow/comments/1wssf4u/ysk_...
https://imgur.com/a/6FaQQhb (original post before being taken down by mods)
>The Work Number is an Equifax (who famously had a massive data breach a few years back) owned database with employment and pay information.
>You can create a login on theworknumber .com and pull your report. Mine is seventy four pages and had info from every company I've worked at in the last 13 years including every paycheck I had received with the exact dollar amount (both net and gross) and hours worked. It has employment start/end dates, termination reason, details about withholdings, benefit enrollment, union affiliation, etc etc etc.
>This data is sourced directly from the HR platform your employees use (ADP, Rippling, Gusto, iSolved, etc). Equifax sells your data for things like employment background checks.
>You cannot have your data removed from The Work Number. You can freeze your report (much like a credit report) but if a potential employer cannot access your frozen report that may disqualify you.
>Why YSK: If you're interviewing for a job you can be at a significant disadvantage especially for things like salary negotiations because they can literally see how much you make including your most recent paycheck. Creditors also can access these reports.
>This is the biggest privacy violation I have ever seen and almost nobody is even aware of it. Certainly none of us consented to having our employment and income data harvested and sold, especially since we get nothing in return. At a minimum, people should know about how this data can impact you.
>Edit: The Work Number is primarily for the US, but there are similar services for other countries.
> In my country (Hungary) you literally get a little pink book where your boss is supposed to sign the time you spent there. You keep it until retirement, the company holds it while you work there. It's also digital now but we're one foot in 1990 and one foot in 2016.
We had something similar in Romania too. Now it's 100% digital and no more delays in informing the government. The employeer needs to notify the government before someone starts work. With the paper version you had a couple of days even weeks if I'm not mistaking.
The other one to check is LexisNexus, which covers auto insurance coverage and traffic violations. In the US, gov't tracked traffic citations fall off after 5 years, but last forever in the baby credit bureaus we're discussing.
Also, INTERPOL. You never know...
FYI https://consumer.risk.lexisnexis.com/request. The Work Number gives you results immediately but LexisNexus says it's going to mail them to you.
Oh, everybody's going to get AI-pentested whether they want to or know about it or not. It's the cost of being on the Internet. Probably the situation will continue to deteriorate. Both the British Library and Jaguar Land Rover recently suffered long outages due to compromises, for example. I suspect we'll probably just lose a few large, famous businesses entirely to compromises.
(Many already do).
That attitude is the problem. Why does our industry have that attitude towards quality? Every bike shop in my little town is better with quality than the average software shop in the world. Yes, software is more complex than bicycle. But a software engineer also gets paid 10x and has the luxury of spending substantial time on their product, compared to the 10 minutes it takes the bike guy down the street to diagnose and then fix an issue with my bike which I then trust my life with once they are done and I bike through traffic.
We need to treat software differently. "Oh well, it's just software shrug" does not cut it anymore, if it even ever did.
You could say that the entire selling point of Apple is that it's better quality than the competitors (even though it's still pretty bad if you compare it to a bicycle).
I guess it's more difficult to spot quality when it comes to software. I would definitely pay for high quality software. It's just that usually the crap quality of a software I'm using only shows slowly over time in little bugs and bad usability and there's no way for me to check for it beforehand.
The owner, and probably user, of a bike is the person paying for the repair. It is to their advantage to ensure it is a good repair.
People expect a repair will be good, and will blame the person who repaired it if not. With software people often blame themselves for issues, and they have no expectation of quality. They also cannot tell quality until after they are committed to using it. Bike shops do not benefit from vendor lock-in.
Personally, I find that we have been flooded with poorly designed, poorly QA'd, break-after-you-use-it-thrice product. I really don't think we have better quality standard in other industry. Quality, testing, design, etc has a cost and most company rather pull out a new version of their product every quarter than actually make a good product. The exception, which goes for software as well, is life critical applications (most public transit, defense, etc).
In ethical terms I agree that software quality should be better, but I think this line is too dismissive. Software is a few orders of magnitude more complicated than a bicycle, to the point where I would say it's an entirely fruitless comparison.
In the simple case, the shop isn't really repairing your bike, it's deploying a fix (a new part) that has already been designed and manufactured. This is a much different and easier thing to do.
In the case of new features, let's say you wanted your bike to also attach a kite, so it went faster on a windy day. The bike shop proposes to do this by attaching the kite to a bolt that they want to put in the frame by drilling a hole all the way through this frame. Is this acceptable to do? Maybe, but who knows, really? How much testing do you want to happen before you trust your life that the frame won't collapse?
There's a lot of effort put into engineering things we take for granted every day. The upside to software is we think we can change it to easily do anything. The downside is if you care about reliability then that ten minute / one line change has hundreds to thousands of hours of test program behind it that we collectively pretend doesn't exist.
When I worked in IT at a hospital, we even had that as an explicit line drawn between 2 departments - IT supported everything that had zero impact on patient health, there was an entirely different department for anything that could impact a patient. And yes, their work was often simple mechanical fixes to health care equipment, while our projects were multi-year reinventions of patient record systems and upgrades to the physical infrastructure of the network, etc. But we never got confused as to which department was actually more important to the patients.
I'm not saying the attitude is OK. It would be lovely if everyone cared about quality. I'm just saying there is a somewhat legit reason for it.
I think we've created a self-fulfilling prophecy. Everyone involved is acting with the best of intentions, but in avoiding what they fear, they've give shape and realized their fears. Much like a greek tragedy.
An example of this is the story of Oedipus Rex, in the story Laius, the king, is told that he is "doomed to perish by the hand of his own son." (and wed his mother) And so to avoid this fate he decides to kill the infant. The person assigned to abandon him in the woods takes pity on the baby and gives the baby away. Thereby ensuring that Oedipus knows neither his mother or his father (and arguably giving him a reason to kill his father).
The child grows up and hears the same prophecy again and the child tries to avoid the prophecy as well, as he loves his adoptive parents. So he leaves them and travels to Laius' kingdom, where he runs into Laius. Neither recognizes the other. As Laius is the type of man to kill an infant, they end up in an argument, whereupon Oedipus kills him.
I think the ancients were on to something, because if Laius had reacted to the prophecy with courage, he would have been saved. I would like to argue that if he had faced his fear and raised Oedipus with love, then the necessary preconditions for the prophecy to come true wouldn't have taken root. But that's not what happens.
By being driven by his neuroses and in acting with cruelty out of fear, Laius makes the prophecy real.
To quote Heraclitus, ethos is fate. Or, character is fate.
I think a lot of people in this AI research sub-culture would be served well by reading these classics, because they are making their self-prophesied doom come true.
They have been convinced for years (GPT-2 was released in Feb 2019) that AI is dangerous. A tremendous threat. An apocalyptic threat.
One dimension of this fear has been the idea that a super smart AI will take over our digital infrastructure and be responsible for the digital apocalypse. That would be terrible!
So what do they do?
They try to make a counter to their fears by teaching models how to exploit vulnerabilities.
How dangerous is such an entity? Very!
Convinced of this danger, they start testing their models as if they were weapons with offensive capability. And then they create models that can be used as weapons.
And because they don't want to release a dangerous weapon out into the world (oh no!), they restrict access to their AI, thereby depriving everyone of tools they can use to improve their security...
Ethos anthropoi daimon.
Much longer than that. Sam Altman said it was an extinction threat before he started Open AI in 2015. It's so dangerous that only he, as the best and greatest examples of humanity, was a good choice to create it, and it was even his responsibility to create it to pre-empt some worse person! A fascinating new variation on the White Mans Burden paradigm.
But this attitude towards dangerousness of AI I simply don't understand.
In my view, AI is obviously risky and dangerous, because it is the only thing on the planet that can think/reason at a human level (or higher) apart from us (and that capability is what made us the uncontested apex species on the planet).
AI is unconstrained by hard biological limits; keeping up with its capabilities will be impossible for baseline humans.
You can argue all day about particulars, like whether selfreplicable robotic shells are required to reach/exceed "existential risk" level to our species (or if artificial minds are enough of a potential threat by themselves).
But the general "risk" is simply AI becoming able to act in its own interests to our detriment, and I don't understand how anyone can dismiss this right now-- help me understand.
Of course there are plenty of imaginable scenarios where coexistence is peaceful and mutually (?) beneficial, but that does not answer those concerns by itself at all.
This year, I've had the UK variously send me someone else's court summons, someone else's tax documentation (not address error - just inadvertently included as attatchments on emails to me).
I've had portugal disclose property taxation information on any subject to me, through a simple enumeration hole.
And last week I had brazil's ministry of agriculture send me details on other peoples' livestock shipments.
Our information environment is unbelievably porous. There would be no secrets from an unshackled AI.
- a UK police force send me a speeding prosecution notice for a car I'd sold 6 months prior to the offence.
- a UK energy company send me to a debt collection agency for unpaid bills they had continued to send to a non-existent address even after telling me they had cleared what was owed and corrected their records of the address.
- a German government official showing up at my door to ask if I was running a specific named business (that I'd never heard of) at the house, which I'd only just taken ownership of from the builder
Someone I used to know in the UK kept getting emails for other people with the same name, including lawyer confidential communications for a namesake who was involved in one of the big banking scandals.
> Our information environment is unbelievably porous. There would be no secrets from an unshackled AI.
Absolutely. My biggest defences against that include the luck of someone else more famous with my own name.
I suspect very few people have had this experience, yes.
AI is a superset of ML, though today most successful AI approaches are based on ML so the line has blurred in casual speech.
Why isn't it obvious? Transformers are deep learning models, and deep learning is a subset of ML.