Computer Science > Computation and Language

arXiv:2101.00288v1 (cs)

[Submitted on 1 Jan 2021 (this version), latest version 1 Jun 2021 (v2)]

Title:Polyjuice: Automated, General-purpose Counterfactual Generation

Authors:Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, Daniel S. Weld

View PDF

Abstract:Counterfactual examples have been shown to be useful for many applications, including calibrating, evaluating, and explaining model decision boundaries. However, previous methods for generating such counterfactual examples have been tightly tailored to a specific application, used a limited range of linguistic patterns, or are hard to scale. We propose to disentangle counterfactual generation from its use cases, i.e., gather general-purpose counterfactuals first, and then select them for specific applications. We frame the automated counterfactual generation as text generation, and finetune GPT-2 into a generator, Polyjuice, which produces fluent and diverse counterfactuals. Our method also allows control over where perturbations happen and what they do. We show Polyjuice supports multiple use cases: by generating diverse counterfactuals for humans to label, Polyjuice helps produce high-quality datasets for model training and evaluation, requiring 40% less human effort. When used to generate explanations, Polyjuice helps augment feature attribution methods to reveal models' erroneous behaviors.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2101.00288 [cs.CL]
	(or arXiv:2101.00288v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2101.00288

Submission history

From: Tongshuang Wu [view email]
[v1] Fri, 1 Jan 2021 18:34:22 UTC (4,313 KB)
[v2] Tue, 1 Jun 2021 17:13:45 UTC (5,708 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2021-01

Change to browse by:

References & Citations

DBLP - CS Bibliography

listing | bibtex

Tongshuang Wu
Marco Túlio Ribeiro
Jeffrey Heer
Daniel S. Weld

export BibTeX citation

Computer Science > Computation and Language

Title:Polyjuice: Automated, General-purpose Counterfactual Generation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Polyjuice: Automated, General-purpose Counterfactual Generation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators