• threeonefour@piefed.ca
    link
    fedilink
    English
    arrow-up
    82
    ·
    3 days ago

    I had to take mandatory AI safety training at work. One thing that they kept stressing was how raw data shouldn’t just be handed over to AI tools.

    If you wanted to analyze customer orders, you should first redact the dataset by removing customer names, addresses, and any other identifiable information before giving it to the AI to ensure company secrets aren’t leaked.

    Is the tool so amazing that it’s worth the risk of exposing trade secrets every time you use it? My company leadership says yes!

    • ivan@piefed.social
      link
      fedilink
      English
      arrow-up
      54
      ·
      3 days ago

      Some companies now have AI trainings, that are AI-generated, like AI-generated avatar with AI-generated voice tells people about AI.

      What in the dystopian cyberpunk hell is that. 🫠

    • MalReynolds@slrpnk.net
      link
      fedilink
      English
      arrow-up
      7
      ·
      3 days ago

      You realize de-anonymizing data is a well researched, effective workflow at this stage? I would be entirely unsurprised at OpenAI and Anthropic using it. Meta are masters at it.

    • uen3@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      2
      ·
      2 days ago

      And the reason this is even tolerated is only because most people don’t understand how LLMs work. Like politicians and their “kill switch” idea if an AI “goes rogure” for example. That there’s like one persistent Claude entity, living in a big room at Anthropic HQ remembering and learning things from all the users it interacts with… Of course it’s not that, it’s an instance of a program that launches and runs the same function every time on whatever inputs it is fed. Sometimes it is set up to feed back old inputs each time so it acts like it has some memory of what your instance previously did. Training a model is sort of like creating a hologram and prompting it is sort of like taking a photo of the hologram from a certain angle.

      They could of course train the next version of the model on everyone’s prompts to the previous version- and probably are. But that’s in no way necessary, and is already sort of a desparation move when they’re running out of input text.

      If the trade secrets / PII you’re foolishly sharing with it are getting leaked that’s not an accidental, inevitable consequence of how AI works. It’s because the company hosting the AI is greedily harvesting and storing all your inputs, for no other reason than you agreed it could in the TOS you didn’t read, and then later oopsie they failed to secure it because they don’t really care and know no one will hold them accountable. Yor prompt is the same as any other text you sent in an input form to any website, the company has it and can decide to only use it for what it has to- or for everything it can get away with.