Authenticity in the Age of Generative AI

1      Introduction

As artificial intelligence (AI) technologies continue to evolve rapidly, driven by an explosion of data, new algorithms, increased computational power, and a swell of investment fueled by hype, generative models—like large language models akin to ChatGPT and image generators—are boasting ever more powerful capabilities. It’s not just the technical know-how of these services that is advancing; the barriers to entry for users are also dropping, making once sci-fi-level technology increasingly accessible.

With this improvement in quality comes a troubling reality: it’s becoming harder for our human senses to distinguish between information produced by our flesh-and-blood counterparts and that which is artificially generated. In fact, solutions that mimic specific styles or even faces are becoming particularly popular.

When we read or, especially, when we view an image or video—assuming it’s not obviously fictional—self-reflection reveals that our brains initially accept what we see as real. We’ve been conditioned this way since childhood; we instinctively trust our perceptions as genuine, even if doubt creeps in moments later due to some flaw we notice.

Generative AI can produce hyper-realistic digital content at such speed and precision that it fundamentally raises questions about credibility in our information-saturated society. Virtually every source of information, aside from reality itself, can be tainted with streams of false data, making it increasingly difficult to ascertain credibility without proper safeguards.

We must brace for the challenges ahead because we are on the verge of a world where content generated by artificial neural networks may outnumber that derived from physical reality or biological entities. Within just five years, our digital landscape could predominantly consist of simulated or generated content.

This writing aims to shed light on the societal and cultural risks lurking in the shadows of generative AI and explores potential tools for reinforcing balance. The study examines not only textual media generated by systems like ChatGPT and other language models but also the problems arising on image, audio, and video-based communication platforms. Among the questions I seek to answer are:

  1. How does generative AI influence the concept and assessment of credibility?
  2. What risks does generative AI pose across different communication platforms?
  3. What technical and societal solutions exist to ensure credibility in the age of generative AI?

We stand on the brink of a potential transformation in how we consume digital content and think critically about it. Are we aware of just how low the threshold for creating lifelike forgeries has become?

2      The Evolution of AI Technologies

The rapid advancement of AI technologies, bolstered by growing popularity, has led to significant breakthroughs in various fields over recent years.

Research aimed at modeling and simulating the human brain dates back to the 1950s, producing foundational concepts and techniques like perceptrons and neural networks. In the 1980s and 1990s, machine learning emerged as a central focus of artificial intelligence research. These machine learning methods enabled computers to learn from data and autonomously create models. Deep learning, which employs multi-layered architectures of neural networks, gained traction in the 2000s, leading to breakthroughs in areas such as image recognition, speech processing, and language comprehension.

2.1     Generative Models (Gen AI)

Generative AI models can produce new, often realistic data based on existing datasets. Generative adversarial networks (GANs) emerged in the 2010s, revolutionizing the field. These networks use two neural networks—a generator and a discriminator—that learn in competition with each other, constantly improving the quality of generated data. The generator’s task is to create data that resembles a specific input dataset, such as a portrait, while the discriminator evaluates the generated data against real data. The generator receives positive feedback when the discriminator cannot distinguish the generated content from its real counterpart. This competitive relationship drives both networks.

In recent years, generative technologies (such as language models) have undergone significant advancements. These models leverage vast amounts of textual data to learn and can generate impressively realistic text. Similarly, image-generating models like DALL-E and Stable Diffusion have played pioneering roles in image creation and manipulation. Today, there are increasingly sophisticated, service-based or open-source models available across various fields.

2.2     The Impact of AI Technologies on Society

AI-based systems are capable of processing and interpreting massive amounts of data, as well as generating significant quantities of content. This relentless production can flood digital channels and carry the risk of unverifiable content. Managing the information deluge and selecting relevant data is becoming a more substantial challenge.

3      The Concept and Significance of Credibility

Credibility is a fundamental value that determines how reliable, authentic, and truthful information or a source is perceived to be.

When doubts arise, assessing and examining credibility can be resource-intensive and costly. In certain instances, it is necessary to verify the origin of the information, the trustworthiness of the source, the context, and whether the information aligns with data from other reliable sources.

3.1     The Importance of Credibility in the Information Society

Credibility and reliability are crucial factors in the functioning of society, especially in the information age. We make decisions based on data we consider credible and trustworthy, whether in work, personal life, politics, or even courtroom proceedings. Thus, it is essential for informed decision-making, trust maintenance, and navigating complex situations.

Credible information enables people to make informed choices. A lack of credibility can lead to misinformation, poor decisions, and distrust, which can negatively impact society—and even democracy.

3.2     Information Overload

The phenomenon of information dilution is not new. It can be intentional or unintentional; in many cases, the actors in the information channels feel it is their duty to provide “interpretation” or explanations, which, even with good intentions, can mislead unsuspecting content consumers.

A crucial change has been that the speed of information dissemination has become virtually instantaneous and accessible to all without filters, leading to a significant increase in risk if not only human actors provide information that is intentionally or unintentionally distorted.

Moreover, generated content is becoming increasingly indistinguishable from real content (it may even appear more authentic), making it harder to determine what we can trust. Authentic, credible sources may gain newfound value for discerning audiences.

4      Platform Risks

Thanks to improved algorithms and stronger hardware, generative AI is becoming more widely available (across multiple media), in greater quantities, with better quality, and at a lower cost, enabling malicious actors to exploit these technologies.

In the past, we could usually distinguish fake content based on image quality, writing style, or audio quality. These indicators were relatively easy to identify, but GAN networks are specifically trained (and RLHF is also aimed at this) to create outputs that are indistinguishable from real content.

4.1     Risks from Text-Based Language Models

4.1.1    Forgery, Spam, and Phishing

Content generated by large language models can facilitate identity theft and phishing through fake emails and messages. Increasingly convincing phishing emails (for example, claiming to be from the Ugandan royal family) or even generated banking websites may appear. Personalized “deepfake” emails can be crafted in the style and tone of real individuals, making it easier for recipients to believe the message is legitimate. The verification processes are underdeveloped, and the situation is exacerbated by the fact that false content spreads faster than authentic content or the ability to recognize and remove it.

With programmable language models, fraudulent websites can be created much more quickly and easily, allowing attackers to harvest personal data or persuade victims to transfer money or share access credentials.

4.1.2    Information Dump

Currently, numerous generated or poorly translated websites occupy prominent positions in popular search results (e.g., https://www.jotudni.hu, https://hirvonal.hu, https://www.unite.ai/hu), primarily aimed at luring unwary visitors to increase ad revenue.

4.1.3    Software Development

Programming, and any area where AI can be applied, will accelerate. We must prepare for more frequent and rapid changes; on the other hand, hackers may also use large language models to uncover malicious codes or security vulnerabilities.

4.2     Generating and Forging Images and Photos

Image-generating models can produce lifelike, fake images (such as Midjourney, Leonardo.ai, Stable Diffusion, Tengr.ai) that can even imitate real people appearing in scenarios that never occurred. This technique can be used for extortion, slander, or political manipulation in extreme cases.

4.3     Audio and Voice Forgery

Opportunities exist to fabricate sounds and speech in recordings or even in real-time. For instance, a malicious actor could manipulate audio recordings or phone conversations using brief voice samples to resemble another person’s voice, deceiving others or voice recognition systems. This poses challenges during fact-checking or in executing fraud and deception.

Singers can be easily imitated, and musical compositions that resemble existing songs can be generated, potentially leading to copyright disputes and complicating the livelihood of artists.

4.4     Video Manipulation Opportunities

As of the preparation of this document, video manipulation isn’t yet a real issue, but the time will inevitably come when video production and manipulation will be as easily executed as with text-based language models. These models will be capable of real-time forgery or manipulation (e.g., swapping individuals in a video call) or even generating entirely new videos (e.g., Pika, Runway ML).

As technology advances, generated videos become increasingly lifelike, making it impossible to distinguish them from real footage. This presents numerous challenges and threats, from questioning video identification systems to “deepfake” videos depicting real individuals in fabricated situations, which can be used for political manipulation, extortion, or slander.

4.5     Virtual Realities and Virtual Companions

Generative AI can also be used to create virtual realities and companions. In the gaming or film industries, for example, the created virtual worlds and characters (NPCs) can become increasingly realistic and detailed.

The content doesn’t have to be pre-programmed; it can adapt to consumer needs. The storyline of a film or series, the traversable areas in a game, or the appearances (looks, voices, styles, etc.) of characters can all be modified, or they might automatically adjust for a more immersive and personalized experience.

Already popular are generated video contents (YouTube channels without faces), personalized virtual companions (e.g., character.ai), which can play various roles ranging from education to adult entertainment (even complemented by live streams), potentially shaking up the structures of the affected industries. New communication channels have transformed our daily interactions or work (mobile phones, chat, remote work), but now it could fundamentally change who, or rather what, is on the other end of the line—leading to significant social shifts. Some predict that 2024 could be the year of robotics…

Today, this doesn’t pose an immediate risk, and we may not even know how we would use it, but metaverses could also be forged, as the thresholds and costs for creating simulations decrease.

4.6     Complex Manipulative Attacks

Generative technologies, by mimicking brain functions, can deceive our minds with high effectiveness. This is particularly true when signals come from multiple directions, constantly and multimodally, leading target individuals to perceive them as more credible.

Malicious attacks can stimulate and simulate not only individuals in critical positions (leaders, doctors, direct connections, family members, etc.) but also groups or even masses. All of this can happen in real-time, personalized, and in response to our reactions, deceiving unsuspecting victims.

4.6.1    Challenges on an Individual Level

In personalized attacks, sophisticated phishing techniques and social engineering methods may be employed. It’s conceivable that messages, calls, or even videos could arrive, seemingly from a close loved one, depicting a kidnapping scenario and appearing credible, with the person apparently requesting that, for instance, an elderly relative send money.

Mass, personalized manipulation could also become achievable with just a few clicks in the near future, whether through even more tailored, innocuous-looking advertisements or recommendation systems. However, their effectiveness and tools could easily raise ethical concerns.

4.6.2    Widespread Attacks

Malicious attacks can stimulate and simulate not only individuals in critical positions (leaders, doctors, direct connections, family members, etc.) but also groups or even masses. This can all happen in real-time, personalized, and in response to our reactions, deceiving unsuspecting victims, especially if signals come from multiple directions, continuously and multimodally, which target individuals may perceive as more credible.

4.6.2.1     Social Media and Recommendation Systems

Algorithms that generate personalized recommendations and content (e.g., Facebook, YouTube feeds) can influence users’ perceptions and behavior, often manipulating and shaping their opinions without them realizing it. These subtle modifications are often difficult to track and verify.

Similarly, mass-generated content can exert a nuanced but definite influence in traditional or social media, enabling malicious actors to manipulate public opinion.

4.6.2.2     Search Engine Manipulation

“SEO hackers” can exploit the operational characteristics of search engines, using targeted mass content creation to manipulate search results. This technique can divert unsuspecting users to fake banking or news sites.

4.6.2.3     Spreading False Information

Using generative AI, entire news portals can be generated, complete with fresh content. However, behind the scenes, virtual agents may disseminate false or not entirely factual news with the intent of disinformation, aiming to manipulate public opinion.

4.7     An Accelerating World

Another important but distinct factor is the potential acceleration of processes in software development and IT. Machines’ performance can be scaled (multiplied or accelerated) at a pace that the human brain is unlikely to keep up with due to foreseeable limitations. We must find ways to adapt to the new circumstances.

5      Technical Solutions for Ensuring Credibility

To address the challenges outlined above, it is vital for researchers, experts, and regulatory authorities to continuously examine the potential dangers of new generative AI technologies and develop protective mechanisms. Additionally, enhancing education and community awareness is crucial for helping people recognize and understand the risks associated with such technologies and critically evaluate the content they consume.

In my view, information providers (be they news portals, search engines, or social media platforms) have a fundamental interest in preserving their credibility while maintaining the freedom of opinion, which may require new tools.

5.1     Applying AI for Credibility Verification

AI may be limited in recognizing forgeries and manipulated content, but given the nature of GAN systems, it is unlikely to offer a definitive solution: the next AI could be trained to bypass the defensive system. Maintenance requires ongoing work, and even then, effectiveness may still be in question.

With this in mind, automated systems could be developed to check algorithms, factual claims, and images against credible sources—or indicate when such resources are unavailable or questionable, requiring caution or further examination (fact-checking AI agents).

Moreover, intelligent search engines supported by AI could assist the work of fact-checkers.

5.2     Verified Sources

Verified, documented sources contribute to the production of quality and reliable content. Both the source of the information (author) and the information itself can be verified. While this practice exists in the scientific community, applying it to everyday matters would be costly in terms of time and resources (manual review).

Practical good examples can be found on the X platform:

  1. multiple verifications for users (email, phone, and even credit card payments are necessary for verified status), and
  2. community moderation, notes (community notes) feature for checking claims made in messages.

This approach could be extended to other platforms, allowing opportunities for cross-platform activities for effective fact-checkers.

A rule and toolset must be developed for source verification, such as a search engine or even agents based on large language models that work with sources considered fundamentally credible.

The most effective approach would be for content providers and search engines to responsibly signal, based on a societal consensus, that high-quality, verified content is present. Similar to server certificates, issuers could be verified not only from a technical or network security perspective but also with an accountable and verifiable certification of credible information sources. Naturally, the opportunity for free expression must be maximally preserved.

5.3     Watermarking (for AI-generated Content)

Watermarking AI-generated content by major providers serves as a tool for labeling generated information. The issue with this approach is not only that it can be circumvented or bypassed (in my view, OpenAI has also abandoned this), but many do not even implement it (e.g., open-source models), thus failing to provide a universal solution, as those attempting to bypass watermarking will always have access to increasingly sophisticated open tools (or even the option to create their own model).

5.4     Digital Signatures (for Marking Original Content)

Marking the authenticity of “original” content could also be an important approach. Optionally, everything where source verification is crucial should be able to be equipped with electronic signatures (from mobile photos to official corporate communications):

  1. We verify with our own signature that we created it, and
  2. possibly with the signature of the device manufacturer (camera sensor, microphone DAC, etc.), who confirms that the information came directly from the sensor after biometric identification (so no digital manipulation occurred).

5.5     Blockchain Technology for Ensuring Credibility

An alternative to digital signatures could be the application of blockchain technology. Blockchain technology provides verifiability and certification, allowing for the validation of content authenticity.

5.6     Verifying the Credibility of Software (Algorithms, Databases, AIs)

Software tools can include algorithms implemented in “traditional” program codes, neural networks, and databases (which may serve as the foundation for AI training or represent the essential weighting—thus the AI itself); as well as the prompts directed at AIs often referred to as Software 2.0.

Each of these can describe the operation of our digital world, making their credibility (ensuring we know what is inside) critical. For each category, we must find ways to ensure security, building a suitable control system around it.

One tool is regular testing, output verification, and the examination of critical elements’ source codes (through verification agents). For instance, Twitter/X published the source code of its recommendation system to demonstrate its credibility.

Internal control systems can convey positive messages, but their ongoing maintenance and external verification are essential.

6      Social Impacts and Responsibilities

To manage the risks posed by generative AI, it is crucial to coordinate the application of technological and societal solutions, as well as to enhance user awareness so that an informed evaluation of the new technology develops, recognizing both positive and negative use cases.

6.1     Human Relationships

Ideally, learning from AI can make us better individuals and improve our coexistence with others (e.g., learning to communicate more effectively). Conversely, relationships with reality may deteriorate or even break if users feel more comfortable in their virtually generated worlds than in their real environments. Generated simulations and virtual worlds can lead to addiction or personality distortions, especially for those predisposed to such tendencies (consider the phenomenon of young men in Japan retreating into isolation).

Over time, ethical questions may arise, akin to “gamification,” if a game, chat, or colorful multimedia virtual platform consistently generates elements that are expected to be most interesting to the visitor, making exit difficult.

6.2     Educating Users and Raising Awareness

The rapid advancement of generative artificial intelligence underscores the need for user education and awareness in media consumption. Consumers need to be sensitized to better recognize potential abuses. Educational programs and initiatives should clarify deepfake technologies, content manipulation techniques, and critical media literacy skills. Such awareness programs should be integrated into educational institutions, while online platforms must create guidelines and informational materials for users.

It’s critical to understand that an image, audio recording, or video can no longer be considered evidence due to the ease of forgery—perhaps only if it was created through a known closed chain, verifiably tracing back to its source. Unfortunately, this can also complicate the proof of genuinely occurred events.

Lawmakers and regulators will need to create new legislation and update existing ones to account for the applications and impacts of generative AI. The laws must address the credibility of generated content, intellectual property rights, and individuals’ rights to protect them from false and misleading information and its consequences.

Given that even a video can easily be forged using deepfake techniques, there will likely be a need for a reevaluation of remote identification and validation processes (e.g., identifying oneself via video).

6.4     The Role of Media and Social Media in Maintaining Credibility

Media institutions and social media platforms must play a key role in preserving credibility and combating misinformation. These platforms must either take responsibility for the content they distribute or proactively work to identify and remove false or manipulated content. Transparency and user involvement in credibility verification are fundamentally important for building trust and disseminating authentic content.

6.5     Labor Market

Without delving into the obvious and often-discussed impacts on the labor market, it’s worth noting the importance of experienced employees who can evaluate credibility and compliance more effectively from the perspective of “credibility.”

7      Closing Thoughts

AI has made a swift and profound impression on the tech sector and could influence the economy and society at a similarly significant scale, albeit at a somewhat slower pace. The extensive use of AI presents opportunities for economic growth, competitiveness, and development while also creating avenues for gray and black markets.

While humanity has experience adapting to technological advancements, further efforts will be needed to tackle the new challenges posed by augmented reality, virtual reality, and generative AI. Future developments, such as multimodal models capable of processing and generating text, audio, images, and video, will open new dimensions in media consumption, communication, and interaction.

After establishing clear ethical standards and laws surrounding AI, developers and users must grasp and accept their responsibilities. It is crucial how technology is applied and for what purposes: the system is not to blame, but rather those who use it.

Where credibility is not critical, such as in entertainment, generative AI opens new possibilities. Control, oversight, and ethical standards remain important in this context, even if the requirements are not as stringent. In other areas, such as medicine and education, where credibility is paramount, defining control points, responsibilities, and legal frameworks is essential for maintaining trust and preventing misinformation.