@david_chisnall@infosec.exchange
I find much easier to explain this technology in terms of virtualized hardware and software executed on it.
The "model" is just a custom special-purpose machine architecture executing the "weights", that are the encoding of the software such machines can run.
The match between the topology of the model and the cardinality of the weigths' matrices is a hint about this simple relationship.
I call the "model" as "vector mapping (virtual) machine" or VMM.
Just like a CPU is built by composing transistors, the fundamental component of a VMM is a "vector reducer", that is a little function from a vector (an array of floats) to a scalar (a float, usually in a carefully crafted range).
Vector reducers have a few nice properties:
they compose well, so you can compute M of them in parallel over the same array of N floats, and you get a map from a N-dimensional vector space to a M-dimensional vector space (a layer, in ANN jergon)they stack well, so you can take the M-dimensional output vector of a layer and feed it to another fleet of O vectors reducers to obtain another O-dimensional vector. Such pipeline effectively map the original M-dimensional vector to a O-dimensional one, appyling a complex non-linear trasform to it.crucially, by carefully picking each layer's parametric function and recording each reducer parameters (usually, polinomial coefficients, aka "weights", bias and threasolds), you can back-propagate statistical errors against your intended output for any given input, iteratively programming the machine.Once you realize we are using the data to iteratively compute a software designed to be executed by a very specific, special purpose machines, everything becomes clear about this technology.
You are not "training" an "artificial intelligence" with IP-infridging datasets. No machine is "learning".
You are just compiling such datasets into an executable, just like you would do with a compiler on C sources.
The datasets are the source code for such compilation process.
And if an x86 binary executed by a CPU retain the source's authors' copyright, the same applies to float executed by a GPU.
Also you stop thinking about such software as a subject: it's never
#ChatGPT who is harming or defaming a person, but
#OpenAI through ChatGPT.
Removing the antropomorphization restores the full chain of accountability.
We should all refuse to talk about "artificial intelligences" that "hallucinate" and speak in terms of defective statistically programmed software that was compiled from illegal and unknown sources and that was likely injected with undetectable backdoors.
We should refuse to call
#MachineLearning what is just a compilation process and to use evocative terms like "latent space" what is just a pattern-preserving statistical heavy-loss compression.
Unfortunately we need to wait for the
#AIbubble to blast, because even
#academia is too dependent on
#BigTech money and subsides to... think clearly about this stuff.
Let's just hope it busts before fascists find a way to leverage either the crash or the tech.
#AI #LLM #security
@Da_Gut@dice.camp @zzt@mas.to @pluralistic@mamot.fr