Abstract Wikipedia talk:Frequently Asked Questions
Add topicAppearance
Latest comment: 1 month ago by Arlo Barnes in topic Correcting common misconceptions, possibly needs a separate page AW:NOT
Correcting common misconceptions, possibly needs a separate page AW:NOT
[edit]Inspired by meta:Requests for comment/The future of Abstract Wikipedia and the comments on it. Here are some misconceptions I've drawn out from the subtext (plus a couple extra I heard from Wikimania participants):
- AW has been going for 6 years and has little to show for it!
- Abstract Wikipedia launched for public contributions on 2026-03-19. Due to the complexity of the project compared to a traditional wikitext wiki, there will be significantly more time before substantial content is on this platform. The bulk of the work is happening at Wikifunctions, where the community is creating and translating new functions. And Wikifunctions itself wasn't able to handle Lexemes until 2024-10-25.
- After a collaboration between Google.org and the AW team in 2022, the outgoing Googlers predicted AW's failure, and it's come true
- The Googlers' criticisms were mostly predicated on the idea that Abstract Wikipedia should be a monolith encompassing the NLG infrastructure as well as all the articles. The WMF decided it was worth building Wikifunctions first, so that it could attract its own community and become a valuable project in its own right.
- The functions used in abstract articles are deprecated / known to be untranslatable in certain languages
- That will continue to happen (see [section re: deprecation]), and it's not a huge problem. Wikidata uses deprecated properties a lot, and other wikis are full of deprecated templates and article message boxes. When a limitation is discovered in an NLG function, it should absolutely be deprecated and replaced with a better one, but the old function can remain in place in the meantime and be useful in the languages where it was working.
- Turning each statement on a Wikidata Item into a sentence will not result in a useful article; why not build a reader mode for Wikidata?
- There's only so much visual flair you can add to a list of RDF triples. If you try building it into a story, you'll have to start introducing NLG aspects, which brings us to Abstract Wikipedia.
- I clicked Special:Random and all I got was two sentences and a picture...
- Yes, most articles in any edition aren't as complete or high-quality as the best articles. Like regular wiki articles, abstract articles can be improved by editing.
- I found an article which is grammatically sound but factually inaccurate! I thought these were backed by Wikidata?
- Wikidata has no guarantee of validity, same as every other Wikimedia project. If there is an issue with Wikidata, you're encouraged to contribute and correct the error! Changes will take a while to propagate through the various caches. It may also be that Wikidata already had the correct information, but it's represented using a combination of qualifiers which the functions don't recognise. Contributions there are welcome too!
- This output is far from fluent English, and in some places it's ungrammatical; how is this useful to enwp?
- First of all, being useful to the English Wikipedia is not one of the project's goals. enwp has a large base of contributors who write many articles in fluent English. This isn't true for all language editions, where even something that would be considered a stub on enwp can be valued because there would otherwise be nothing at all. And finally, when the output in English (or any other language) is stilted, you can go to Wikifunctions and fix it!
- If abstract articles don't produce good English, the output in other languages must be far worse
- At the moment this is a sound generalisation. Many of the functions to render abstract content in natural language are first configured in English. Once a function is configured in any language, the quality of prose depends on the quality and complexity of the implementation in that language.
- Some articles may be worse than others, but until they're all good, any Wikipedia would object to having AW integration enabled
- No Wikimedia project is perfect, all suffering from low quality content to some degree. Abstract Wikipedia articles work on an opt-in basis, meaning that communities, not the WMF or Abstract Wikipedia editors, choose which articles to include in their wikis. For an example of this system, see this special page on the Wikipedia Test wiki. 'Integration' in this context means that the special page is enabled on your wiki.
Editors will also be able to create a local article using the abstract article as a starting point, once integration is enabled, or they could use individual functions in an existing article via embedded function calls.
- No Wikimedia project is perfect, all suffering from low quality content to some degree. Abstract Wikipedia articles work on an opt-in basis, meaning that communities, not the WMF or Abstract Wikipedia editors, choose which articles to include in their wikis. For an example of this system, see this special page on the Wikipedia Test wiki. 'Integration' in this context means that the special page is enabled on your wiki.
- An abstract article can't be mapped to every language because there are huge differences in the way languages divide sentences into parts of speech (syntax)
- Human languages can be very different (even without considering "constructed" languages), but at their core they are methods of communication, where an utterance/sentence is a message which encodes some information that the speaker/writer wants to convey. This is the 'semantic' part: the objects and concepts being referred to and the relationships between them. (There are a few concepts which are hard to translate, but not impossible. See [section re: untranslatability].) The rest, the 'syntax' and 'grammar', is where the differences are.
Abstract Wikipedia's content is "abstract" in the sense that it contains only the semantic information of a sentence; in order to actually produce text, it relies on a complex set of NLG algorithms from Wikifunctions. These algorithms are each dedicated to a particular language/dialect in order for them to best handle the grammatical rules of that language.
- Human languages can be very different (even without considering "constructed" languages), but at their core they are methods of communication, where an utterance/sentence is a message which encodes some information that the speaker/writer wants to convey. This is the 'semantic' part: the objects and concepts being referred to and the relationships between them. (There are a few concepts which are hard to translate, but not impossible. See [section re: untranslatability].) The rest, the 'syntax' and 'grammar', is where the differences are.
- Other NLG projects have tried and failed to solve the problem of when to include "the" (the definite article)
- This idea seems to have stemmed from a misunderstanding of an anecdote in Wikifunctions' newsletter, where one of GF's developers is said to have called the definite article problem
one of the hardest puzzles they had to solve
. That issue of the newsletter set out the problem as it applied to WF, and a heuristic solution was provided within a day. That solution is of course imperfect, but perfection was never the goal of Abstract Wikipedia.
- This idea seems to have stemmed from a misunderstanding of an anecdote in Wikifunctions' newsletter, where one of GF's developers is said to have called the definite article problem
- Some concepts are simply untranslatable; languages are in some way limited by culture/worldview
- This is an idea linguists call "linguistic relativity", though historically the debate has been over whether worldview is limited by language. Both that conjecture and the reverse turn out to be false in most aspects. On the idea of 'untranslatability', enwp says
difficulty of translation does not always carry deep linguistic relativity implications; denotation [i.e. identification of concepts by words] can virtually always be translated, given enough circumlocution
. For example, the label of the Wikidata property 'video' in Tyap is 'ghwughwu a̱guguut' which means something like "movement of shadows".[TODO have someone confirm that, or find a better example]
- This is an idea linguists call "linguistic relativity", though historically the debate has been over whether worldview is limited by language. Both that conjecture and the reverse turn out to be false in most aspects. On the idea of 'untranslatability', enwp says
- AW is trying to create an IAL / universal language
- While IALs or areal auxlangs could be target languages, the LISP-like syntax does not itself constitute an auxiliary language. Natural language text is considered the end product of work on an abstract article.
- The foundation of AW is a model of linguistics which is discredited by the current consensus of linguists
- Abstract Wikipedia as a platform does not assume any particular model or theory. All the linguistic information and processing is stored in Wikidata and Wikifunctions, so even if that were the case, the foundations could be swapped out and the project salvaged. This idea may have arisen from [section re: definite article]. See also [section re: known bad functions].
- By accepting contributions, AW leeches volunteer man-hours from sister projects
- We volunteers own the time we spend. If we consider it best spent here, we are obviously aligned with the aims and viability of this project. Closing it would not encourage us to contribute more. Some of us have contributed to Wikimedia projects for decades, and the variety of projects or contributions is one of the things that can help us to motivate and maintain such commitment. Abstract Wikipedia relies on the content in Wikifunctions and Wikidata to produce output, so it pushes contributors to those projects.
- Volunteer time would be better spent using Extension:ContentTranslation
- Since article integration hasn't been rolled out yet, this is definitely true for a lone editor with the goal of making knowledge accessible to more people, and it will likely be true for a while. But if all goes well, Abstract Wikipedia (or rather, Wikifunctions) will accrue the functions needed to generate text in many languages. Then that lone editor will have the choice between translating a high-quality article into one language, or converting some of it into an abstract article which is then available in many languages.
- The server resources would be better spent on training LLMs for under-resourced languages
- Other organisations already do this, so WMF-funded machine translation development would be duplicative. Also, it isn't in the competency area of the Wikimedia movement, which is volunteer-led content. The suggestion to "just" train LLMs for all the under-resourced languages doesn't account for the scale of the problem either: At time of writing, ~250 languages are offered in Google Translate, while there are several thousand living languages.
- It was already possible to create formulaic articles from Wikidata using templates and/or modules, an entirely new system is unnecessary
- Certainly it was possible, some Wikipedias did build systems which failed and others built systems which worked and are still in use. Like the current Abstract Wikipedia, these had a high learning curve for those looking to contribute to the system, but unlike Abstract Wikipedia they are all limited to their own wikis.
- If it succeeds, AW will bring about a sort of "digital neo-colonialism" in which small language communities get enwp translated into their language
- Abstract Wikipedia is a blank slate, not based on enwp. It uses (via Wikifunctions) a multilingual labelling system similar to Wikidata, enabling volunteers to translate the interface which other volunteers can then use to edit articles. We will continue working at lowering the barriers to entry, so that readers of small wikis can feel comfortable clicking
Edit abstract
and sharing the knowledge which is most important to them.
On the subject of electronic colonialism: This is a significant problem for the Wikimedia movement as a whole, not just Abstract Wikipedia. For example, for our language codes we use w:ISO 639, originally developed by SIL Global (an organisation derived from the evangelical missionary group Summer Institute of Linguistics which has a goal of translating Christian Bibles into many languages[relevance?]). By foregrounding the volunteer editors of a target language and enabling self-governance of the individual wikis, these sociopolitical influences may be mitigated.
- Abstract Wikipedia is a blank slate, not based on enwp. It uses (via Wikifunctions) a multilingual labelling system similar to Wikidata, enabling volunteers to translate the interface which other volunteers can then use to edit articles. We will continue working at lowering the barriers to entry, so that readers of small wikis can feel comfortable clicking
- When a new set of NLG functions emerges to fix the flaws of the previous set, every abstract article will need to be rewritten
- Yes, that process is what's known as a 'deprecation cycle' and it's a part of any serious software project. Other wikis will have encountered something similar, for example the migration to Parsoid and the maintenance tasks it created.
Regarding flawed NLG functions specifically: The old function(s) can remain in place and be useful in the languages where they were working, until a replacement exists and an editor swaps it out. And we can be more confident that the replacement(s) will last, because they will have been born from community discussion and fortified with a growing corpus of test cases for catching regressions.
- Yes, that process is what's known as a 'deprecation cycle' and it's a part of any serious software project. Other wikis will have encountered something similar, for example the migration to Parsoid and the maintenance tasks it created.
- Even the most fluent machine-generated text will remain soulless compared to a human-written article, so why bother?
- The point of writing an encyclopedia is conveying factual information, and for Abstract Wikipedia, specifically in the languages where human resources are scarce; the quality of prose is a separate concern, addressed at abstract:WikiProject Encyclopedic Quality. The option of having a human write the whole article will always be there, see [section on article integration].
- Some languages have a lot of inflected forms of words; just storing them on Wikidata is hard, to say nothing of using them in Wikifunctions
- Those languages with many inflections tend to also have consistent rules for how to generate them (which makes sense when you remember that humans find rote memorisation hard, but learning patterns easy). Programmers can use those inflection rules to "compress" the list.
- AW is an automated/machine translation service, something that everyone else is now using LLMs for
- Abstract Wikipedia does not translate between languages. It uses community-defined algorithms to transform an abstract representation of concepts into rich text in a particular written language (NLG).
The reason why AW uses functions it that it gives users complete control over how content is displayed and translated. Whereas LLMs can be unpredictable, hard to properly diagnose issues, and very difficult for volunteers to fix.
- Abstract Wikipedia does not translate between languages. It uses community-defined algorithms to transform an abstract representation of concepts into rich text in a particular written language (NLG).
- Reading AW requires running huge JavaScript blobs in your browser
- When articles are eventually selected for display in language editions of Wikipedia, the reading experience will be similar to any other page on that wiki. The AW site is primarily designed for content development. The editor is currently a 0.5 MiB download. The site itself is also under development, so the current state does not necessarily represent a long-term ideal.
YoshiRulz (talk) 02:25, 28 July 2026 (UTC)
- Thanks, good summary. Should we prepare answers individually here (separate sections?) then add them to the page when done? --99of9 (talk) 02:51, 28 July 2026 (UTC)
- You're welcome to chop up by above comment for organisational purposes. YoshiRulz (talk) 03:00, 28 July 2026 (UTC)
- Okay, I'll write as though it's the answer, anyone should feel free to overwrite and improve. --99of9 (talk) 03:29, 28 July 2026 (UTC)
- You're welcome to chop up by above comment for organisational purposes. YoshiRulz (talk) 03:00, 28 July 2026 (UTC)
- re: assignment of language codes, I can see how that process was Eurocentric, but I don't see what negative impact it's had on those communities who didn't get a memorable code. They're 2- and 3-letter identifiers used mainly by machines. YoshiRulz (talk) 22:42, 28 July 2026 (UTC)
- Well, it depends on what you call 'negative'. The main controversies have been about lumping and splitting:
zhencompasses a wide range of languages, whereasnbandnnare mutually-intelligible varieties of the same language. It doesn't end up being that big a deal since as you say mostly humans don't need to encounter the codes all the time, and there are ways to append the codes to make them more specific. And while WMF bases codes on ISO 639, there are custom decisions likebe-tarask. Meanwhile, Abstract Wikipedia uses Wikifunctions identifiers, which are yet a third system. — Arlo Barnes (talk) 22:47, 28 July 2026 (UTC)
- Well, it depends on what you call 'negative'. The main controversies have been about lumping and splitting:
See also f:WF:NOT. — Arlo Barnes (talk) 20:30, 28 July 2026 (UTC)